AI Research Scientist, Learning & Evaluation

StudyFetchBeverly Hills, CA
$180,000 - $280,000Onsite

About The Position

We are a technology company building AI-native learning products used by over seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both platforms are powered by the Learn Engine, the intelligence that guides a learner from initial understanding to demonstrated mastery. We collaborate with partners like NVIDIA to deliver responsible, learning-first AI to students in need. This role is crucial for establishing how we measure the effectiveness of AI tutors, going beyond public benchmarks to understand true student comprehension, model behavior, and long-term retention. You will develop the evaluation and measurement framework that underpins model training, product decisions, and learning science across both StudyFetch and Honen. As a founding team member, you will work directly with decision-makers and set the standards for all models and features.

Requirements

  • PhD in statistics, computer science, machine learning, economics, physics, computational social science, or another quantitative field, plus 5+ years applying it to real products or research, OR a track record of owning evaluation or measurement for an AI product that shipped to real users.
  • Experience evaluating LLMs in production, including designing eval suites, identifying successes and failures, and proposing improvements.
  • Strong understanding of causality, experimental design, controls, sample size, and the limitations of observational data.
  • Clear and effective written and verbal communication skills, capable of presenting to diverse audiences including engineers, founders, and external researchers.
  • Genuine curiosity about AI and its applications.
  • A strong commitment to the mission of improving learning outcomes for students.
  • Proficiency in Python (Pandas, NumPy, SciPy, Jupyter).
  • Expert-level SQL skills.
  • Experience with large-scale datasets.
  • Knowledge of experimental design, causal inference, Bayesian and frequentist methods, and hypothesis testing.
  • Familiarity with AI/LLM concepts such as eval frameworks, LLM-as-judge, RAG, embeddings, agent workflows, and fine-tuning.
  • Experience with data technologies like MongoDB, PostgreSQL, vector databases, and warehouse/pipeline tooling.
  • Comfort with cloud infrastructure (GCP) and reasoning about inference cost, latency, throughput, and GPU utilization.
  • Experience building dashboards and reporting tools.

Nice To Haves

  • Learning science, psychometrics, or item response theory.
  • Experience working with children's data and associated regulations.

Responsibilities

  • Develop and implement evaluation strategies for AI models, assessing accuracy, reasoning, child safety, resistance to sycophancy, and Socratic teaching methods.
  • Design, run, and maintain internal benchmarks for multi-turn tutoring conversations, defining measurement criteria and defending choices to external researchers.
  • Translate product data into training signals and evidence, analyzing which interventions impact mastery versus engagement and correlating model behaviors with actual student learning.
  • Measure agreement between subject-matter experts for content verification, identify outdated answer keys, and design feedback loops for training data.
  • Build dashboards and reporting for product and leadership teams to track analytics such as retention, activation, feature adoption, and conversion rates, and their correlation with model changes.
  • Identify and work with engineering to implement missing events, telemetry, and logging for currently unanswerable questions.
  • Establish the measurement standards for the team, including experiment execution, defining results, and determining when changes should be shipped.

Benefits

  • $180,000–$280,000 base salary, plus equity
  • 100% employer-paid Medical, Dental, and Vision
  • 75% dependent coverage for Medical, Dental, and Vision
  • 401(k) with employer matching
  • Daily team dinner provided in-office
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service