Staff Machine Learning Engineer

ServiceNowSanta Clara, CA
Hybrid

About The Position

The Agentic Engineering organization at ServiceNow is the customer-obsessed engineering group that builds a conversational AI experience that turns enterprise intent into completed work. We advance how enterprise AI reasons, remembers, and executes. The Agent Orchestration team — the team you'll join — owns the execution core: the agent harness, orchestration runtime, multi-agent coordination, memory management, and the evaluation frameworks that ensure agents behave correctly in production. Every autonomous action Otto promises depends on what this team ships. By joining our team, you’ll be at the forefront of our AI transformation journey, backed by the global scale of ServiceNow and the agility of a high-growth environment. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. You will design, build, and operate production-grade agentic AI systems embedded across ServiceNow's platform — autonomous agents that reason over real enterprise data, take action across workflows, and run safely at Fortune 500 scale.

Requirements

  • 6+ years building production software systems with a strong track record on reliability, performance, and scalability
  • Hands-on experience shipping generative AI products — not just integrating LLM APIs or building prototypes, but owning AI-powered features that production users depend on
  • Solid depth in how large language models work: failure modes, context constraints, and how prompt design shapes model behavior at scale
  • Practical prompt engineering experience: systematically designing, versioning, and evaluating prompts across model updates or A/B evaluation cycles
  • A real track record in eval engineering — not just familiarity, but a portfolio of evaluation suites designed, shipped, and used to drive quality decisions in production AI systems
  • Cost and efficiency awareness at the system level: experience reasoning about model routing, inference cost, and latency tradeoffs in production
  • Strong software engineering fundamentals: distributed systems, API design, and testing discipline
  • Comfort operating in fast-moving, ambiguous, startup-like AI product environments

Nice To Haves

  • Exposure to AI services deployed in a micro services-based architecture (Kubernetes, OpenShift, etc.)
  • Experience tracing, debugging, identifying root causes, and resolving issues in AI codebases that span multiple services, multiple environments, and/or multiple tenants
  • Published work, patents, or open-source contributions in Machine Learning, AI, distributed systems, etc.

Responsibilities

  • Design and ship multi-agent systems — orchestration, tool use, planning loops, memory, and failure recovery — that operate reliably in production, not in notebooks.
  • Build agents that leverage ServiceNow's data layer — CMDB, Workflow Data Fabric, and Knowledge Graph — to make decisions with context no frontier model has on its own.
  • Own the guardrails: observability, human-in-the-loop controls, and compliance infrastructure that make autonomous systems safe to deploy at scale.
  • Work closely with our search team to ensure agents are grounded in accurate, low-latency retrieval — RAG pipelines, hybrid search, re-ranking, and evaluation — as a critical dependency of agentic quality.
  • Integrate frontier models (Anthropic, Google, OpenAI) into the Sense → Decide → Act → Govern architecture; evaluate trade-offs across cost, latency, and capability for production use cases.
  • Raise the technical bar through architecture decisions, code reviews, and coaching — particularly on agentic design patterns and production AI discipline.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service