Head of Evaluations Research

AaruNew York, NY
Onsite

About The Position

Evaluation Research determines whether Aaru's populations, predictions, and simulations correspond closely enough to the real world to support consequential decisions. The function defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. As Head of Evaluation Research, you will set Aaru's evaluation charter and lead the research needed to test its core systems. The strongest evaluations will be grounded in observed behavior and outcomes, including transactions, product usage, operational records, resolved events, and longitudinal decisions. You will work across Population Research, Prediction Research, Simulation Engineering, and customer-facing teams while preserving the independence needed to identify inconvenient results. This is a hands-on research leadership role. You will design studies, construct evaluation datasets, write analysis code, inspect individual failures, and develop new measurements when existing benchmarks are inadequate. You will also recruit and lead a small team that can combine scientific rigor with a practical understanding of how research and engineering systems improve.

Requirements

  • Developed an original evaluation or measurement agenda in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or an environment with a comparable bar for rigor.
  • Built evaluations that changed a research direction, model capability, product decision, or scientific conclusion.
  • Can define a difficult construct precisely enough to measure it without losing the underlying question.
  • Comfortable with experimental design, observational data, sampling, uncertainty, statistical power, leakage, and condition shift.
  • Can write code, analyze large datasets, design studies, and inspect individual model failures.
  • Can work closely with the teams building a system while reaching independent conclusions about its quality.
  • Care more about an accurate result than a favorable one and are willing to revise your own evaluation when it proves inadequate.
  • Can explain technical evidence clearly to researchers, engineers, customers, company leadership, and the public.
  • Led researchers or a major technical direction while remaining directly involved in the work.
  • Want to build in person, in New York, at high speed.

Nice To Haves

  • Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, or measurement theory.
  • Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, or decision-support tools.
  • Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or validation against operational outcomes.
  • Experience building evaluation platforms, regression suites, experiment-tracking systems, or shared research datasets.
  • A record of finding an important failure that standard metrics missed and developing a better way to measure it.
  • Experience communicating scientific or technical results in customer-facing, public, policy, or regulatory settings.

Responsibilities

  • Define a coherent evaluation agenda across population construction, predictive systems, agent behavior, group dynamics, and end-to-end simulations.
  • Turn broad questions about realism, accuracy, and usefulness into measurable constructs, decisive experiments, and clear decision criteria.
  • Build tests of population quality that assess whether each generated person forms a coherent whole, whether the population reproduces important relationships in the data, and whether rare but plausible profiles are represented.
  • Evaluate forecasts and other predictive outputs using temporal holdouts, prospective outcomes, calibration, ranking quality, subgroup performance, and the real cost of different errors.
  • Compare simulations with transactions, behavioral traces, product adoption, operational outcomes, market movements, and other records of what people actually did.
  • Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another.
  • Develop end-to-end studies that show whether better components lead to better answers for the decisions customers use Aaru to make.
  • Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, or changing environments.
  • Establish strong baselines, clean holdouts, contamination controls, and statistical standards appropriate to each research question.
  • Create diagnostic evaluations that help researchers identify why a system failed and whether a proposed fix generalizes.
  • Work with Simulation Engineering to make evaluations repeatable and versioned, and convert production outcomes and customer failures into durable test cases.
  • Produce evidence for customers and the public that is reproducible, appropriately scoped, and clear about uncertainty and limitations.
  • Communicate negative and inconclusive results with the same care as positive findings.
  • Hire, mentor, and lead exceptional evaluation researchers and research engineers while remaining a direct contributor.

Benefits

  • Aaru has a clear and widely trusted definition of quality across populations, predictions, and end-to-end simulations.
  • Core evaluations are grounded in observed outcomes and reveal whether performance generalizes across time, domains, and groups.
  • Researchers receive measurements that are diagnostic enough to guide improvement while protected holdouts preserve the integrity of final results.
  • Important failures are found early, explained clearly, and converted into durable tests.
  • Product and company decisions rely on evidence that is reproducible, appropriately uncertain, and connected to real-world value.
  • Customers and the public can understand what Aaru has demonstrated, where the evidence is limited, and how confidence should be interpreted.
  • A small, exceptional team develops a reputation for evaluation work that is scientifically rigorous, practically useful, and unusually honest.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service