About The Position

NVIDIA is seeking an Evaluation/ML-Systems Engineer to join their AI Safety & Security Engineering team. This role is critical in measuring the program's effectiveness, ensuring that real capability is distinguished from anecdote. The engineer will be responsible for defining benchmarks, designing metrics, choosing baselines, and writing protocols for fair comparisons. They will build infrastructure to ensure reproducibility and traceability of results, automating measurement processes to allow researchers to focus on core questions. The role involves partnering with security researchers and platform engineers to design trustworthy and repeatable experiments, and fostering a culture of evidence-based conclusions within the team.

Requirements

  • Bachelor's degree (or equivalent experience) with 5+ years in ML engineering or evaluation.
  • Evaluation experience: Designing benchmarks, metrics, and statistically sound comparisons for ML systems.
  • Measurement rigor: A careful, skeptical approach to metrics, baselines, and claims.
  • Engineering skills: Solid Python engineering for shared infrastructure, including experiment tracking and data pipelines.

Nice To Haves

  • Security evaluation: Exposure to evaluating security tooling or pipelines.
  • Agentic systems: Experience measuring agent or LLM behavior.
  • Community work: Contributions to public benchmarks or evaluation frameworks.

Responsibilities

  • Build the benchmarking and reproducibility systems.
  • Define the metrics and protocols for measurement.
  • Map every result to the code and runs that produced it.
  • Keep findings reviewable and conclusions traceable.

Benefits

  • Competitive salaries
  • Generous benefits package
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service