Staff Machine Learning Engineer

ServiceNowSanta Clara, CA

About The Position

We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms. This isn't a thin LLM wrapper. We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our from-scratch, industry-leading agent execution runtime). Our systems are backed by a first-party MCP surface into admin/product/rules/transaction systems and interoperate with other AI agents over the A2A protocol. Below that sits a conventional Java/Spring Boot microservices fleet and a React/TypeScript + Lit frontend that the agents ultimately drive. We're looking for someone who already operates at a Senior-Staff bar in the agentic/LLM domain but is building out breadth across the rest of the stack. You'll be one of the most senior technical voices on how agentic systems get designed here — state management, tool boundaries, streaming protocols, prompt/context architecture, and multi-agent coordination — while staying credible end-to-end: able to read a Spring Boot service, unblock a frontend integration, or reason about a classical ML model pipeline when the problem calls for it.

Requirements

  • 8+ years building production software, with several years specifically shipping LLM-powered / agentic systems (not just API wrapper calls — real tool-use loops, state management, multi-turn orchestration)
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
  • Deep, hands-on expertise with LangGraph and/or LangChain (or the judgment to know when to skip them and hand-roll something better)
  • Strong understanding of MCP — ideally having built an MCP server, not just consumed one
  • Production async Python (FastAPI, asyncio) — comfortable with WebSockets, streaming, and the failure modes of long-lived stateful connections
  • Track record of making real architecture decisions on agent systems — tool boundaries, context/state design, cost/latency tradeoffs, when a bounded agent beats a fully autonomous one
  • Enough range outside Python to read/modify a Spring Boot service and a React component without hand-holding — this is explicitly a whole-stack role, not "Python specialist with a frontend allergy"
  • Comfort operating with ambiguity and setting technical direction, not just executing a spec — this is a Staff-level bar on judgment
  • A demonstrated habit of pulling the latest from the industry — new agent frameworks, protocol standards, model capabilities — into production quickly
  • Willingness and ability to build real fluency in the Core ServiceNow platform.
  • Willingness to work directly with customers/deployments as part of Forward Deployed Engineering efforts

Nice To Haves

  • Experience with classical ML (PyTorch/scikit-learn) in addition to LLM-based systems
  • Experience with A2A or other agent-to-agent interop protocols
  • Experience with RAG / knowledge-graph systems (embeddings, vector or graph-based retrieval)
  • Multi-tenant SaaS experience, especially around per-tenant isolation of stateful connections/resources
  • CPQ, quoting, or transaction/pricing domain experience
  • Prior Forward Deployed Engineering experience, or time spent embedded with customers shipping bespoke solutions exposure

Responsibilities

  • Design and implement multi-agent orchestration systems using LangGraph/LangChain agents over frontier LLMs for transaction editing, conversational configuration, and multi-product quote planning.
  • Develop and own the evolution of 'The Harness', our proprietary agent execution runtime, focusing on full-duplex sessions, low-latency streaming, backpressure, live progress, partial results, and graceful cancellation.
  • Influence the strategic direction of MCP as a secondary interface versus the Harness as the primary agent interaction mechanism.
  • Contribute to the A2A protocol for agent-to-agent task delegation and streaming.
  • Engage in Forward Deployed Engineering, working with customer- and product-facing teams against live deployments to adapt the Harness and agents to real-world workflows and constraints.
  • Develop and refine RAG / context engineering strategies, including tenant-uploaded document ingestion, categorization, and aggregation into agent context, and prefix-cacheable prompt design for cost/latency.
  • Evaluate and potentially extend classical ML pipelines (PyTorch/scikit-learn) for problems not suited for LLMs, and compare them against LLM-based alternatives.
  • Demonstrate full-stack fluency by reading/modifying Spring Boot/Java services and React components to unblock end-to-end integrations.
  • Operate with ambiguity and set technical direction for agent systems, making architectural decisions on tool boundaries, context/state design, cost/latency tradeoffs, and agent autonomy.
  • Stay current with industry advancements in agent frameworks, protocol standards, and model capabilities, and integrate them into production systems.
  • Build fluency in the Core ServiceNow platform to enable direct interoperation with its systems.
  • Participate in Forward Deployed Engineering efforts, working directly with customers and deployments.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service