Senior Staff (DevOps) Machine Learning Engineer

ServiceNow•Santa Clara, CA
•Hybrid

About The Position

The AI Services Platform Group at ServiceNow is a customer-focused innovation group building intelligent software and smart user experiences using existing and latest advanced technologies to enable end-to-end, industry-leading work experiences for customers. We are a group of researchers, applied scientists, engineers, and product managers with a dual mission. We build and evolve the AI platform, and partner with teams to build products and end-to-end AI-powered work experiences. In equal measure, we lay the foundations, research, experiment, and de-risk AI technologies that unlock new work experiences in the future. We are seeking a Senior Staff Machine Learning Engineer (DevOps) to lead the architecture, optimization, and operational excellence of our AI inference platform. You will own the end-to-end inference lifecycle—from onboarding cutting-edge models to deploying optimized inference engines in production across multiple backends on Kubernetes infrastructure. You'll also architect the intelligent gateway layer using Envoy and ext_proc to enforce access control, provide observability, and enable advanced request routing. This role combines deep technical expertise in GPU compute, LLM APIs, inference frameworks, containerization, orchestration, and API gateway architecture with a bias toward pragmatic engineering and measurable performance impact.

Requirements

  • Expertise in Go or Python, OOP, Design Patterns, time and space-efficient algorithms
  • Experience building new products that use challenging algorithms
  • Expertise in coding efficient, object-oriented, modularized and quality software
  • Knowledge of core AI/ML techniques and algorithms
  • Knowledge of unit testing, profiling, and code tuning
  • Deploy, scale, and operate services in production Kubernetes environments, including Helm-based deployments and CI/CD pipelines that make releases repeatable and safe to roll back.
  • Own observability for what you build — meaningful metrics, useful logs, real tracing, and alerts that fire on customer impact rather than on noise.
  • Debug and resolve production incidents independently, participate in on-call, run root-cause analysis, and operate against defined service level objectives.
  • Artificial intelligence strategy is a critical skill at Experienced proficiency for IC5 — it applies to how you build and to what you build.
  • Confident use of AI coding assistants such as GitHub Copilot or Cursor, paired with the critical eye to review AI-generated code as rigorously as any other — catching logic errors, security anti-patterns, and missed edge cases rather than treating model output as production-ready.
  • AI-assisted debugging and trace analysis, and a clear understanding of the data privacy rules governing what goes into a prompt.
  • Mandatory minimum 8+ years of Backend software engineering experience
  • 4+ years of designing and operating distributed systems
  • 3+ years of Kubernetes in production, including Helm and CI/CD
  • Bachelor’s degree in computer science, software engineering, or a closely related technical field required.
  • A master’s degree may offset up to one year of the experience minimum.
  • Equivalent practical experience is considered where technical depth is clearly demonstrable.

Nice To Haves

  • GPU compute
  • LLM APIs
  • inference frameworks
  • containerization
  • orchestration
  • API gateway architecture

Responsibilities

  • Own the end-to-end inference lifecycle—from onboarding cutting-edge models to deploying optimized inference engines in production across multiple backends on Kubernetes infrastructure.
  • Architect the intelligent gateway layer using Envoy and ext_proc to enforce access control, provide observability, and enable advanced request routing.
  • Deploy, scale, and operate services in production Kubernetes environments, including Helm-based deployments and CI/CD pipelines that make releases repeatable and safe to roll back.
  • Own observability for what you build — meaningful metrics, useful logs, real tracing, and alerts that fire on customer impact rather than on noise.
  • Debug and resolve production incidents independently, participate in on-call, run root-cause analysis, and operate against defined service level objectives.
  • Apply artificial intelligence strategy to how you build and to what you build.
  • Confidently use AI coding assistants such as GitHub Copilot or Cursor, paired with the critical eye to review AI-generated code as rigorously as any other — catching logic errors, security anti-patterns, and missed edge cases rather than treating model output as production-ready.
  • Utilize AI-assisted debugging and trace analysis, and have a clear understanding of the data privacy rules governing what goes into a prompt.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service