Senior Software Engineer, Agent Eval Platform

ServiceNowMountain View, CA
$161,300 - $274,200Onsite

About The Position

Moveworks' AI agents are designed to act, plan, call tools, and modify enterprise systems on behalf of millions of employees. The core challenge for this role is to develop a robust measurement system for evaluating agent performance across multi-step interactions. This involves building the 'judgement layer' of the agent evaluation platform, including rubrics, judges, calibration against human labels, and methodologies to ensure scores are meaningful and can be used to train agents to improve. This role focuses on applied ML in a challenging area where LLMs are being used to judge other LLMs, particularly in complex enterprise environments with real-world actions. The position is not focused on pretraining or traditional testing but on developing novel methodologies for evaluating agentic AI. The role offers opportunities to contribute to three key areas: Eval orchestration at scale, Agent observability and tracing, and Stateful simulation. In Eval orchestration, responsibilities include building a runtime for multi-turn agent scenarios, managing scheduling, retries, and execution isolation, versioning specs and datasets, and establishing reliability and SLOs for the evaluation harness. For Agent observability and tracing, the focus is on adopting OpenTelemetry, defining a span data model for agent trajectories, ensuring trace context propagation, maintaining full prompts and completions, and enabling fault attribution. In Stateful simulation, the work involves creating simulation environments with stateful fakes of enterprise systems, managing data injection and teardown for hermetic runs, developing LLM-driven user simulators, and contract-testing mocks. Across all areas, the goal is to lay the groundwork for using evaluation signals to optimize agent performance.

Requirements

  • Experience in at least 3 of the following areas: Distributed systems (idempotency, delivery guarantees, isolation, determinism, reproducibility), Orchestration and workflow runtimes (DAG execution, scheduling, retries, backfills, high-concurrency job systems like Temporal, Airflow, Argo), Observability internals (OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data), Concurrent and async programming (Python asyncio, Go concurrency, structured cancellation), Data-intensive pipelines (high-volume ingest, schema evolution, sampling and retention trade-offs), gRPC/protobuf service and interface design.
  • 5+ years building production backend or infrastructure systems.
  • Strong proficiency in Python or Go (ideally both).
  • Experience designing and operating systems that handle real traffic at scale.
  • Comfort making non-deterministic systems measurable.
  • Ability to turn fuzzy agent behavior into a signal engineers are willing to gate releases on.
  • Comfort with ambiguity and novel problems without textbook solutions.

Nice To Haves

  • An ML background is not required, but an interest in turning fuzzy agent behavior into a measurable signal is beneficial.

Responsibilities

  • Build the judgement layer of the agent evaluation platform, including rubrics, judges, and calibration against human labels.
  • Develop methodologies to score agent actions precisely enough to teach them to improve.
  • Create a runtime that executes multi-turn agent scenarios end-to-end, managing the user-agent-world loop, collecting data, and running validators.
  • Implement scheduling, retries, high-concurrency execution, and run isolation for production dataset sizes.
  • Establish versioned specs, datasets, and reports with run-to-run comparison capabilities.
  • Consolidate existing eval workflows onto a single orchestration service.
  • Establish a reliability floor and SLO for the evaluation harness.
  • Enable self-serve evaluation capabilities for any team.
  • Lead the adoption of OpenTelemetry-native observability for the agent platform.
  • Define the span data model for agent trajectories, making them queryable.
  • Ensure trace context propagation across asynchronous boundaries and long-lived sessions.
  • Maintain full prompts and completions throughout the pipeline and prevent eval traffic contamination.
  • Develop fault attribution and cross-run diffing capabilities.
  • Support the debug surface for harness engineers and define the tracing contract with agent development teams.
  • Build the simulation environment, including stateful fakes of enterprise systems.
  • Implement per-run data injection and programmatic setup/teardown for hermetic and repeatable runs.
  • Develop LLM-driven user simulators and scripted state-machine simulators.
  • Implement contract-testing mocks against real API schemas in CI.
  • Lay the foundation for using eval signal to optimize the agent, not just measure it.

Benefits

  • Health plans
  • Flexible spending accounts
  • 401(k) Plan with company match
  • ESPP
  • Matching donations
  • Flexible time away plan
  • Family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service