Member of Technical Staff, Enterprise Evals Platform

MercorSan Francisco, CA
$220,000 - $425,000Onsite

About The Position

Mercor is an AI data company building the layer between human expertise and frontier models. The Enterprise agents are complex systems that require reliable and economically viable work. Evaluation is key to achieving this, encompassing checking correctness and optimizing routing based on cost, latency, and quality. This role involves decomposing real work, capturing expert standards, and encoding them to prevent agents from shortcutting tasks. The successful candidate will apply Mercor's learnings from building benchmarks with domain experts to devise new methods for improving evals, rubrics, and the agents measured against them. This is a platform engineering role with a focus on evals, requiring the building of verifiers, agent measurement environments, and scalable grading infrastructure that abstracts across customers, domains, and tasks.

Requirements

  • Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails.
  • Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed.
  • Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it.
  • Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems.

Nice To Haves

  • Experience with Harbor environments and RL environments.

Responsibilities

  • Define golden sets: decompose real tasks and encode the expert quality bar.
  • Build verifiers over agent trajectories and outputs, calibrated and hard to game.
  • Build the eval platform that runs offline environments, task suites, and grading at scale.
  • Run loss analysis over production trajectories and turn failure modes into regression tests.
  • Run the optimization loop across models, prompts, skills, and harnesses.
  • Own the rollout gates that decide whether an agent change ships.
  • Partner with the Enterprise Platform team and the Applied AI engineers embedded with customers.

Benefits

  • Up to $15k relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Generous equity grant vested over 4 years
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service