AI Evaluation Infrastructure Engineer

Block•Bay Area, CA, United States of America, CA
•$263,600 - $395,400•Remote

About The Position

We build AI products, and the quality of our evaluations sets the ceiling for how good those products can be. The speed of our evaluations determines how quickly we can improve them. We are looking for an engineer to build the infrastructure and tooling that make high-quality AI evaluation possible at Block's scale. You will help teams understand whether a model or product change is actually better, whether a result is statistically meaningful, and whether offline evaluation is predicting what happens with real users. Our evaluation approach combines offline evals that encode our definition of a good response, online evals that show how people actually respond, and a feedback loop that keeps the two converging. Your work will turn that approach into systems that product teams can use quickly, reliably, and with confidence. This is a high-impact, early-stage area with broad surface area. You will help decide what to build first, then build the platform that helps teams ship better AI products faster.

Requirements

  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.
  • Strong statistical literacy, including comfort with confidence intervals, variance, power, and multiple comparisons.
  • The judgment to identify when a result is meaningful and when it is noise.
  • Experience evaluating LLM or ML systems, or deep systems engineering experience with a strong interest in AI evaluation.
  • Product instinct for internal tools. You understand that leaderboards, annotation tools, and workflows only matter if teams actually use them.
  • A bias toward building reliable, observable systems that other engineers can trust.
  • Strong collaboration skills and the ability to work across ambiguous product, data, and engineering problems.

Responsibilities

  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.
  • Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.
  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.
  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.
  • Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance, so teams can distinguish real improvements from noise.
  • Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.
  • Build the loop that compares offline scores with online outcomes, identifies eval sets that have stopped predicting reality, and helps teams improve them.
  • Partner with product, engineering, data, and ML teams to make evaluation workflows fast enough and trustworthy enough to become part of everyday development.

Benefits

  • Remote work
  • medical insurance
  • flexible time off
  • retirement savings plans
  • modern family planning
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service