AI Research Scientist, Learning & Evaluation

StudyFetchBeverly Hills, CA
Onsite

About The Position

We are a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most. This role is crucial for defining and measuring what 'better' means in AI tutoring, ensuring that our models are trained and tuned for actual learning outcomes rather than just performance on benchmarks. You will build the evaluation and measurement layer that underpins model training, product decisions, and learning science across both StudyFetch and Honen. As a founding-team role, you will work directly with decision-makers and set the standards for all models and features.

Requirements

  • PhD in statistics, computer science, machine learning, economics, physics, computational social science, or another quantitative field, plus 5+ years applying it to real products or research, OR fewer credentials with a track record of owning evaluation or measurement for an AI product shipped to real users.
  • Experience evaluating LLMs in production.
  • Ability to clearly articulate the design, outcomes, and lessons learned from an evaluation suite.
  • Rigorous understanding of causality, experimental design, controls, and sample size.
  • Ability to differentiate between a real effect and a dashboard anomaly.
  • Clear and concise written and verbal communication skills.
  • Proficiency in Python (Pandas, NumPy, SciPy, Jupyter).
  • Expert-level SQL skills.
  • Experience with large-scale datasets.
  • Knowledge of experimental design, causal inference, Bayesian and frequentist methods, and hypothesis testing.
  • Familiarity with AI/LLM evaluation frameworks, LLM-as-judge and its failure modes, RAG, embeddings, agent workflows, fine-tuning, and post-training.
  • Experience with data tools such as MongoDB, PostgreSQL, vector databases, and warehouse/pipeline tooling.
  • Comfort with cloud platforms like GCP and reasoning about inference cost, latency, throughput, and GPU utilization.
  • Experience building dashboards and reporting tools.

Nice To Haves

  • Learning science, psychometrics, or item response theory.
  • Experience working with children's data and associated regulations.

Responsibilities

  • Develop and own the evaluation framework for AI models, covering accuracy, reasoning, child safety, resistance to sycophancy, and Socratic teaching methods.
  • Design, run, and analyze evaluations for every candidate model, holding the release bar for new checkpoints.
  • Build and maintain an internal benchmark for multi-turn tutoring conversations, defining its metrics and potentially publishing parts of it.
  • Translate product data into training signals and evidence, identifying which interventions improve mastery versus engagement, and correlating model behaviors with student learning.
  • Measure agreement between subject-matter experts, identify outdated answer keys, and design feedback loops for content verification.
  • Develop dashboards and reporting for product and leadership on key analytics such as retention, activation, feature adoption, and conversion, and their relation to model changes.
  • Identify missing events, telemetry, and logging, and collaborate with engineering to implement them to enable new insights.
  • Define the team's measurement standards for experiments, results, and feature shipping.
  • Present findings and insights to engineers, founders, and external research partners.
  • State uncertainty quantitatively when possible and qualitatively otherwise.

Benefits

  • $180,000–$280,000 base salary, plus equity
  • 100% employer-paid Medical, Dental, and Vision; 75% dependent coverage
  • 401(k) with employer matching
  • Daily team dinner provided in-office
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service