Staff AI Platform Engineer, Infrastructure Services

SentinelOne
$156,000 - $215,000

About The Position

As a Staff AI Platform Engineer, Infrastructure Services, you will be tasked with taking ownership of our AI Gateway infrastructure (built on Kong AI Gateway), the system that authenticates, routes, rate-limits, and monitors AI coding assistant traffic org-wide, while also being fluent enough across our broader platform stack to design solutions that span the two. This is a high-autonomy, high-scope role: you will set technical direction for AI infrastructure, drive incident response and reliability work, and partner closely with the engineers who own our CI/CD, GitOps, and artifact systems rather than working in isolation from them.

Requirements

  • 8 or more years of experience in platform, infrastructure, or DevOps engineering, with a track record of owning systems end-to-end in production.
  • Hands-on experience with API gateway technologies (Kong, Envoy, Apigee, or similar); direct experience with AI/LLM gateway patterns (rate limiting, semantic caching, prompt/response observability) is a strong plus.
  • Strong Kubernetes and GitOps experience (ArgoCD or comparable), and comfort operating across multiple environments (dev, gov, prod).
  • Solid CI/CD background: Jenkins pipeline design and administration, build infrastructure, and runner/agent fleet management (GitHub Actions runners or equivalent).
  • Experience with artifact and package management systems (Artifactory, Xray, or similar) and source control platform administration (GitHub Enterprise).
  • Working knowledge of infrastructure-as-code (Terraform) and cloud platforms (AWS/EKS).
  • Experience deploying and operating self-hosted LLM inference stacks (vLLM, NVIDIA Triton/NIM, TGI, Ollama, or similar) and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
  • Familiarity with LLMOps practices: model versioning, evaluation harnesses, and usage/cost observability across API-based and self-hosted models.
  • Track record of setting technical direction, driving cross-team initiatives, and mentoring other engineers; this role has significant scope and minimal day-to-day oversight.
  • Clear, proactive communicator who can explain infrastructure trade-offs to both engineers and non-technical stakeholders.

Nice To Haves

  • Experience operating LLM/AI-assisted developer tooling at scale (Claude Code, Copilot, or similar) inside an enterprise is preferred.
  • Familiarity with Okta/OIDC and enterprise auth patterns for internal platforms is preferred.
  • Experience with engineering productivity metrics tooling (LinearB or similar) and AI-based code review tooling (Qodo or similar) is preferred.
  • Experience with vector databases and RAG pipelines (e.g. Milvus, Pinecone, pgvector, or similar) in a production setting is preferred.
  • Exposure to model fine-tuning or lightweight training pipelines (LoRA/QLoRA or similar) for domain-specific model adaptation is preferred.

Responsibilities

  • Work on the AI Gateway platform: architect, harden, and scale our Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS), including auth (Okta/OIDC), consumer tiers and budgets, rate limiting, semantic caching, and observability.
  • Lead reliability and incident response: drive root-cause analysis and remediation for gateway issues (timeouts, latency, capacity, failover) and build the monitoring/alerting needed to catch them before users do.
  • Design across the platform, not just the gateway: work fluently with our CI/CD (Jenkins, JPAAS), GitOps and Kubernetes deployment tooling (ArgoCD across dev/gov/prod), artifact management (Artifactory/Xray), GitHub Enterprise administration, and GitHub Actions runner fleet, so that AI infrastructure decisions account for how the rest of the platform actually works.
  • Evaluate and roll out AI developer tooling: run structured pilots and adoption efforts for tools like AI-assisted PR review (Qodo) and engineering metrics platforms (LinearB), and make clear build-vs-buy recommendations.
  • Set technical direction and mentor: define architecture and standards for AI infrastructure, review designs across the team, and raise the bar for other engineers working in this space.
  • Partner cross-functionally: work directly with security, DevEx, and product engineering teams consuming the gateway to translate their needs into platform capabilities.
  • Host and serve local models: stand up and operate self-hosted/open-weight model serving infrastructure (e.g. vLLM, NVIDIA Triton/NIM, TGI, Ollama) for workloads where routing to an external provider isn't the right fit, including GPU capacity planning, autoscaling, and cost/performance tuning.
  • Support the broader model lifecycle: help build LLMOps practices such as model versioning, evaluation, and safe rollout, plus supporting infrastructure for retrieval-augmented generation (vector stores, embedding pipelines) as use cases mature.
  • Track usage and cost: build observability into token usage, latency, and spend across both API-based and self-hosted models so the business can see what AI infrastructure actually costs.

Benefits

  • Equity & Rewards
  • Restricted Stock Units (RSUs)
  • Employee Stock Purchase Plan (ESPP)
  • Time Off & Wellbeing
  • Flexible time off
  • Paid company holidays and paid sick time
  • Gender-neutral parental leave
  • Grandparent leave
  • Insurance & Financial Security
  • Medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Life and disability insurance
  • Health and dependent care FSA
  • Voluntary benefits (hospital, accident, critical illness)
  • Employee Assistance Program (EAP)
  • ARAG pre-paid legal
  • Nationwide pet insurance
  • Cancer Care program
  • Global business travel medical insurance
  • Work Perks & Flexibility
  • Home office allowance
  • Mobile phone reimbursement
  • Wellness & Lifestyle
  • Wellness coach
  • Wellness/gym reimbursement
  • Fertility coverage
  • Adoption & surrogacy reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service