Sr. DevOps Engineer II, AI Platforms

DoubleVerify•New York, NY
•$153,000 - $260,000•Hybrid

About The Position

Over the past year, DV built an AI platform at speed. More than 800 people now use Claude and Cursor daily against an MCP gateway that fronts 30+ internal tool backends, a gateway for model traffic, a registry where engineers publish skills and plugins, and a tracing pipeline running off every laptop. It was built by a strong, opinionated group of engineers who continue to drive it forward, and it is already in wide production use. This is the first role at DV dedicated entirely to that platform. You'll join the engineers building it today with AI infrastructure as your full focus, which means having the time to go deep on the hardest problems and keep the platform moving day to day. A lot of it is already live and needs operating and hardening. Plenty is still ahead of us: group-based authorization across the gateway, a consolidated multi-region model gateway with real failover, and cost attribution that teams can act on.

Requirements

  • 5+ years in DevOps, Platform Engineering or SRE, with production Kubernetes ownership and the software engineering habits to go with it.
  • Hands-on experience managing and scaling AI infrastructure in production: Demonstrated track record operating developer AI tools, model routing/gateway platforms, or LLM observability stacks at organizational scale.
  • Real depth in OAuth 2.0/2.1, OIDC and JWT: token exchange, scope design, claim propagation, and federation between an IdP, an authorization server and a proxy. Some of the hardest problems on this platform live here.
  • Production experience with a gateway, proxy or service mesh (Envoy or equivalent), and the judgment to debug a misbehaving control plane under pressure.
  • A track record of taking infrastructure from prototype to GA. You've set the SLOs, built the failover, written the runbooks and carried the pager.
  • The collaboration skills to work well inside an opinionated, distributed engineering group. You build on other people's work, disagree constructively, and write things down.

Nice To Haves

  • Upstream contributions to Envoy Gateway or a comparable open-source proxy.
  • LLM observability tooling (Langfuse, LangSmith, Arize Phoenix) and OTEL GenAI conventions.
  • Cost attribution or chargeback for a shared platform.
  • Keycloak or equivalent IdP engineering.
  • Zero-trust access platforms and secrets management.

Responsibilities

  • Run and harden the MCP gateway day to day. It runs on Envoy Gateway on GKE, secured by Keycloak, serving 30+ backends to Claude Code, Cursor and Claude Desktop across the company, plus an externally reachable gateway. You'll be one of the first responders when it breaks.
  • Take the lead on closing the authorization gap: group-based access from JWT claims, per-tool RBAC, and scoped OAuth consent in place of the static user lists and broad scopes we run on today.
  • Help consolidate the model gateway estate onto one supported deployment and drive it to GA, with multi-region, health-checked failover and coverage beyond a single client.
  • Push AI observability to GA with the team. Finish the trace pipeline, extend it across tools, and give teams per-session, per-user and per-department cost and latency numbers they can trust.
  • Build out the infrastructure behind AI cost governance: spend caps that reflect real dollars rather than raw tokens, plus attribution, enforcement and alerting.
  • Support the registry, eval and marketplace platform that lets engineers publish and consume skills, plugins and MCP servers safely, and partner with Security on model and data access control, prompt injection exposure and connector supply chain review.
  • Work across a DevOps org spread over three regions, pairing with the engineers who built each piece.

Benefits

  • bonus/commission (as applicable)
  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service