Moveworks' AI agents are designed to act, plan, call tools, and modify enterprise systems on behalf of millions of employees. The core challenge for this role is to develop a robust measurement system for evaluating agent performance across multi-step interactions. This involves building the 'judgement layer' of the agent evaluation platform, including rubrics, judges, calibration against human labels, and methodologies to ensure scores are meaningful and can be used to train agents to improve. This role focuses on applied ML in a challenging area where LLMs are being used to judge other LLMs, particularly in complex enterprise environments with real-world actions. The position is not focused on pretraining or traditional testing but on developing novel methodologies for evaluating agentic AI. The role offers opportunities to contribute to three key areas: Eval orchestration at scale, Agent observability and tracing, and Stateful simulation. In Eval orchestration, responsibilities include building a runtime for multi-turn agent scenarios, managing scheduling, retries, and execution isolation, versioning specs and datasets, and establishing reliability and SLOs for the evaluation harness. For Agent observability and tracing, the focus is on adopting OpenTelemetry, defining a span data model for agent trajectories, ensuring trace context propagation, maintaining full prompts and completions, and enabling fault attribution. In Stateful simulation, the work involves creating simulation environments with stateful fakes of enterprise systems, managing data injection and teardown for hermetic runs, developing LLM-driven user simulators, and contract-testing mocks. Across all areas, the goal is to lay the groundwork for using evaluation signals to optimize agent performance.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed