Applied AI Researcher, Agent Systems & Evaluation

NuroMountain View, CA
Hybrid

About The Position

Nuro is seeking an Applied AI Researcher specializing in Agent Systems & Evaluation to join their Frontier Models team. This role focuses on building and rigorously evaluating AI agents that operate autonomously within Nuro's engineering organization. The goal is to amplify engineer and researcher output by 100x through a trustworthy and evidence-based approach to AI development. The team operates with a startup mentality, leveraging Nuro's decade of experience in self-driving technology to create a robust closed-loop evaluation system for AI work. This position involves hands-on model work, including fine-tuning and RL on open-source vision-language models using proprietary driving data, and requires a deep understanding of frontier models, their evaluation, and their real-world impact.

Requirements

  • Graduate degree in CS, ML, statistics, or a related field, or equivalent research experience.
  • Demonstrated research judgment.
  • Fluent in the current literature and able to judge it.
  • Deep understanding of how LLMs work — pretraining through the post-training stack, and what actually happens at inference.
  • Reason from mechanism, not just from published numbers.
  • Firm grasp of the full evaluation pipeline: sourcing eval data, constructing the loop, automating the climb.
  • Having done all three for a real system, rather than one in isolation.
  • Rigorous experimentalist: design experiments that can fail, understand variance and power, and be comfortable saying an intervention didn't work.
  • Hands-on post-training experience — SFT and RL, ideally on open-weight models — including the data curation and evaluation work required to know whether it actually helped.
  • A real engineering background.
  • Strong Python, comfortable with production systems and data, able to stand up the infrastructure your own experiment needs.
  • Direct experience with LLM agent systems — building them, evaluating them, or studying why they fail.
  • Measure yourself in impact and in weeks.
  • Want your work in front of hundreds of engineers this quarter.

Nice To Haves

  • Published or applied work in agent evaluation, reasoning, test-time compute, RL, or verification.
  • Experience with multimodal or vision-language models, and with data curation at scale.
  • Online experimentation in production: A/B testing, causal inference from observational data, offline-to-online correlation.
  • Experience building evaluation harnesses, task suites, or LLM-as-judge systems, including their failure modes.
  • Familiarity with autonomous systems, safety cases, or verification-gated deployment.
  • Vision-language model experience is a strong plus.

Responsibilities

  • Make design decisions about agent systems based on evidence rather than intuition.
  • Test hypotheses against genuine production traffic from the first month.
  • Improve the agent system, which is complex enough that intuition is no longer sufficient.
  • Ensure the system performs against real-world data, including Nuro's codebase, infrastructure, and engineers' actual requests.
  • Turn any task into a closed loop: define success, identify evaluation data sources, collect signal, and feed results into the next iteration.
  • Own the evaluation pipeline end-to-end, including eval data collection, eval loop construction, and automated hill climbing.
  • Mine production traces for labeled outcomes and capture human accept/reject/edit signals.
  • Build task sets that reflect the real distribution of work.
  • Determine when a model-based judge is trustworthy.
  • Turn fuzzy objectives into measurements that run on every change.
  • Design experiments for production settings where clean randomization is not always available.
  • Search for improvements through prompts, context strategies, tool sets, routing, reasoning budgets, and model choice once a task has a trustworthy loop.
  • Post-train models on proprietary driving data using supervised fine-tuning and RL on open-source vision-language models.
  • Specify needed data using a labeling workforce.
  • Read the frontier research and convert it into experiments within weeks.
  • Map where additional inference compute pays and where a verifier beats a bigger model.
  • Work in close partnership with the engineer who builds the platform, determining what the system should be doing and whether it worked.

Benefits

  • Annual performance bonus
  • Equity
  • Competitive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service