Evaluation Research Manager

AaruNew York, NY
$280,000 - $425,000Onsite

About The Position

Aaru builds simulations of human behavior using AI agents to model real-world populations and decision-making. Companies use these simulations to test choices before committing to product launches, pricing, communications, or policy changes. Aaru emphasizes the need for simulations to represent real people, be calibrated, remain coherent, and present evidence legibly. The company is a small, in-person team in New York, valuing urgency, high ownership, and intellectual honesty, expecting team members to surface inconvenient evidence, adapt quickly, and see work through to completion. Evaluation Research at Aaru focuses on ensuring that the company's populations, predictions, and simulations accurately reflect the real world to support critical decisions. This team defines measurement standards, develops methodologies, and generates evidence for system improvement and capability description. The function operates as both a creator of reusable evaluation infrastructure ('rails') such as datasets, harnesses, and reporting systems, and a performer of specific evaluations ('carts') like backtests, forecast studies, and coherence tests. It is distinct from conventional QA or internal approval, functioning as an independent research unit that collaborates with development teams while maintaining the ability to communicate potentially unfavorable findings.

Requirements

  • Led evaluation, measurement, or empirical research in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or a comparably rigorous environment.
  • Built evaluations that changed a research direction, model capability, product decision, or scientific conclusion.
  • Can define a difficult construct precisely enough to measure it without reducing away the underlying question.
  • Comfortable with experimental design, observational data, sampling, statistical power, uncertainty, causal threats, leakage, and condition shift.
  • Can write code, analyze large datasets, design studies, inspect individual failures, and review the technical work of researchers and engineers.
  • Can work closely with builders while reaching independent conclusions about the quality of their systems.
  • Care more about an accurate result than a favorable one and are willing to revise your own evaluation when evidence shows it is inadequate.
  • Can prioritize a research portfolio and choose which uncertainty is most important to resolve next.
  • Have managed or technically led strong researchers, give clear feedback, and can develop independent judgment rather than creating dependence on your review.
  • Can explain technical evidence clearly to researchers, engineers, product teams, customers, company leadership, and non-specialists.
  • Want to work in person in New York with a team that moves quickly and takes truth-seeking seriously.

Nice To Haves

  • Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, survey methodology, or measurement theory.
  • Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools.
  • Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, backtesting, or validation against operational outcomes.
  • Experience building evaluation platforms, regression suites, experiment-tracking systems, shared research datasets, model scorecards, or scientific reporting tools.
  • A record of finding an important failure that standard metrics missed and developing a better measurement method.
  • Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings.
  • Experience hiring and leading a small, high-talent research team through ambiguous work with short feedback cycles.

Responsibilities

  • Build, lead, and develop a high-performing team of Evaluation Researchers and research engineers.
  • Translate broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, decisive experiments, and explicit decision criteria.
  • Set a focused evaluation agenda across population construction, predictive systems, individual agent behavior, group dynamics, and end-to-end simulations.
  • Decide which evaluation infrastructure should become a reusable organizational rail and which questions require a purpose-built study.
  • Establish standards for baselines, temporal holdouts, prospective testing, contamination control, statistical power, uncertainty, subgroup analysis, and reproducibility.
  • Build tests of population quality that assess individual coherence, joint and conditional distributions, representation of rare but plausible profiles, and whether a profile induces behavior consistent with the person it represents.
  • Evaluate forecasts and other predictive outputs using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the real cost of different errors.
  • Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, and longitudinal decisions.
  • Design end-to-end studies that determine whether improvements to a component actually improve the decision-relevant output customers receive.
  • Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, changing environments, or ambiguous ground truth.
  • Create diagnostic evaluations that help researchers localize why a system failed and distinguish a real general improvement from benchmark-specific optimization.
  • Partner with Simulation Engineering to make evaluations repeatable, versioned, scalable, and integrated into development and release workflows without compromising protected holdouts.
  • Convert production incidents, customer surprises, and deployment failures into durable test cases and better measurement methods.
  • Review evidence used in product, customer, or public claims and ensure that conclusions are reproducible, appropriately scoped, and honest about uncertainty and limits.
  • Communicate negative, null, and inconclusive results with the same precision and urgency as positive findings.
  • Recruit exceptional researchers, set clear expectations, provide direct feedback, develop independent research judgment, and address performance problems early.

Benefits

  • competitive base salary
  • equity participation
  • comprehensive medical, vision, and dental coverage
  • visa sponsorship and relocation support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service