Mercor is an AI data company building the layer between human expertise and frontier models. The Enterprise agents are complex systems that require reliable and economically viable work. Evaluation is key to achieving this, encompassing checking correctness and optimizing routing based on cost, latency, and quality. This role involves decomposing real work, capturing expert standards, and encoding them to prevent agents from shortcutting tasks. The successful candidate will apply Mercor's learnings from building benchmarks with domain experts to devise new methods for improving evals, rubrics, and the agents measured against them. This is a platform engineering role with a focus on evals, requiring the building of verifiers, agent measurement environments, and scalable grading infrastructure that abstracts across customers, domains, and tasks.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed