Senior Software Engineer, Agent Eval Platform

ServiceNowMountain View, CA
Hybrid

About The Position

Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing. This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

Requirements

  • Experience in at least 3 of these: Distributed systems: idempotency, delivery guarantees, isolation, and — unusually central here — determinism and reproducibility. Orchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high-concurrency job systems (Temporal, Airflow, Argo, or something you built yourself). Observability internals as a builder, not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data. Concurrent and async programming: Python asyncio, Go concurrency, structured cancellation. Data-intensive pipelines: high-volume ingest, schema evolution, sampling and retention trade-offs. gRPC/protobuf service and interface design.
  • 5+ years building production backend or infrastructure systems
  • Strong in Python or Go (ideally both)
  • Experience designing and operating systems that handle real traffic at scale
  • Comfort making a non-deterministic system measurable. You don't need an ML background — but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases on
  • Comfort with ambiguity; these are novel problems without textbook solutions

Responsibilities

  • Eval orchestration at scale: The runtime that executes multi-turn agent scenarios end-to-end — stand up the environment and user simulator, drive the user↔agent↔world loop, collect transcripts, traces, and final state, run validators and scoring, tear down. Scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes. Versioned specs, datasets, and reports, with run-to-run comparison as a first-class operation. Consolidating evals that run today as one-off workflows onto a single orchestration service — one source of truth, one place to schedule and retry. Establishing a reliability floor and an SLO for the harness itself. Getting to self-serve, so any team runs an eval without bespoke integration.
  • Agent observability and tracing: Leading the move to OpenTelemetry-native observability for the agent platform, replacing the parallel per-service logging, correlation, and redaction mechanisms in use today. The span data model for agent trajectories — prompts, tool calls, plan updates, outcomes — so a trajectory is queryable, not reconstructed by hand from log files. Trace context propagation across async boundaries and sessions that stay alive for minutes or hours. Making full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own data. Fault attribution and cross-run diffing: which component actually broke, and what changed since the last green run. The debug surface support and harness engineers use, and the tracing contract with the team that builds the agent.
  • Stateful simulation: The simulation environment itself: stateful fakes of the enterprise systems agents call — ITSM, HR, knowledge bases, inventory — backed by a real datastore that persists changes during a run, so a created ticket is visible to a later read. Per-run data injection and programmatic setup/teardown so every run is hermetic and repeatable. LLM-driven user simulators for open-ended personas, and scripted state-machine simulators for deterministic flows. Contract-testing mocks against real API schemas in CI, so simulation fidelity can't quietly drift as vendor APIs change. Ahead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads — provisioned from an identical baseline and torn down every run.
  • Across all three: laying the foundation for using eval signal to optimize the agent, not just measure it.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service