Applied AI Researcher, Agent Systems & Evaluation

NuroMountain View, CA
$193,930 - $352,290

About The Position

Nuro is seeking an Applied AI Researcher specializing in Agent Systems & Evaluation to join their team. This role focuses on building and rigorously evaluating AI agents that operate autonomously within Nuro's engineering organization. The goal is to amplify engineer and researcher output by 100x by ensuring autonomous work is trustworthy and reliable. The position involves developing a closed-loop evaluation system for AI work, leveraging frontier transformer models, and adapting them with proprietary data. The researcher will own the end-to-end evaluation pipeline, from data collection to automated hill climbing and post-training model work. This role requires a strong engineering background, research judgment, and a focus on real-world impact and scaled outcomes. The researcher will work closely with platform engineers and leadership, with significant autonomy in decision-making.

Requirements

  • Engineering background with strong research taste.
  • Demonstrated research judgment.
  • Fluent in the current literature and able to judge it.
  • Deep understanding of how LLMs work (pretraining through post-training, inference).
  • Reason from mechanism, not just published numbers.
  • Firm grasp of the full evaluation pipeline: sourcing eval data, constructing the loop, automating the climb.
  • Experience with all three stages of the evaluation pipeline for a real system.
  • Rigorous experimentalist: design experiments that can fail, understand variance and power.
  • Comfortable stating when an intervention didn't work.
  • Hands-on post-training experience (SFT and RL), including data curation and evaluation.
  • Strong Python skills.
  • Comfortable with production systems and data.
  • Able to stand up the infrastructure for experiments.
  • Direct experience with LLM agent systems (building, evaluating, or studying failures).
  • Measure self in impact and weeks; desire to have work in front of hundreds of engineers.
  • Graduate degree in CS, ML, statistics, or related field, or equivalent research experience.

Nice To Haves

  • Published or applied work in agent evaluation, reasoning, test-time compute, RL, or verification.
  • Experience with multimodal or vision-language models.
  • Experience with data curation at scale.
  • Online experimentation in production (A/B testing, causal inference from observational data, offline-to-online correlation).
  • Experience building evaluation harnesses, task suites, or LLM-as-judge systems, including their failure modes.
  • Familiarity with autonomous systems, safety cases, or verification-gated deployment.
  • Vision-language model experience is a strong plus.

Responsibilities

  • Make system design decisions based on evidence rather than intuition.
  • Test hypotheses against genuine production traffic.
  • Improve agent systems where intuition is no longer sufficient.
  • Extract maximum value from existing frontier models and adapt them with proprietary data.
  • Ensure agent systems perform against real-world data, including Nuro's codebase, infrastructure, and engineers' actual requests.
  • Turn any task into a closed loop, defining success, evaluation data, signal collection, and iteration.
  • Own the evaluation pipeline end-to-end, including eval data collection, eval loop construction, and automated hill climbing.
  • Mine production traces for labeled outcomes and capture human accept/reject/edit signals.
  • Build task sets reflecting the real distribution of work.
  • Determine when a model-based judge is trustworthy.
  • Turn fuzzy objectives into measurements that run on every change, considering noise floors and statistical standards.
  • Design experiments for production settings where clean randomization is not always available.
  • Optimize prompts, context strategies, tool sets, routing, reasoning budgets, and model choice.
  • Post-train models on proprietary driving data using supervised fine-tuning and RL on open-source vision-language models.
  • Specify needed data using a labeling workforce.
  • Read the frontier research and convert it into experiments, separating replicable results and turning them into live experiments.
  • Focus on test-time scaling, including reasoning budgets, sampling and search strategies, verifier-guided selection, and escalation decisions.
  • Map where additional inference compute pays off and where a verifier beats a bigger model.
  • Work in close partnership with the platform engineer to determine what the agent system should be doing and whether it worked.

Benefits

  • Annual performance bonus
  • Equity
  • Competitive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service