About The Position

At NVIDIA, we redefine what’s possible in digital imaging, personal computer gaming, and high-performance computing. Now, we lead the charge into AI’s unlimited potential. Our GPUs act as the brains of computers, robots, and self-driving cars. They enable these machines to understand and interact with the world in new ways. This is your chance to join a legacy of innovation and excellence. Work with the world’s best talent in a diverse and encouraging environment. Join us and make a lasting impact on the world! This is an ambitious opportunity to work at the forefront of AI safety and contribute to groundbreaking advancements in technology.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, data science, or a related technical field.
  • 10+ years of experience in technical program management, engineering, product development, technical operations, or a similar area.
  • Strong understanding of LLM architecture, frameworks (e.g., OpenAI, Anthropic, Hugging Face), and model evaluation.
  • Familiarity with LLM development, post-training, inference, tool calling, evaluation datasets, and model-release lifecycles.
  • Experience directing programs related to AI/ML development, model evaluation, agentic safety, security, content safety, hallucinations, and/or production releases.
  • Familiarity with AI safety risks such as hallucinations, timely injection attacks, unsafe tool use, data poisoning, model manipulation, and unintended agent behavior.
  • Experience supporting model-safety evaluations, Red Teaming, adversarial testing, security assessments, or responsible-AI programs.
  • Ability to interpret technical evaluation findings and communicate their product, schedule, and release implications to leadership.

Nice To Haves

  • Experience with LLM Agentic Safety, Security, Red Teaming, Hallucinations, adversarial assessment, or Responsible AI.
  • Experience defining evaluation datasets, rubrics, benchmarks, and release gates.
  • Ability to use evaluation results and data to make clear release recommendations and communicate residual risk to leadership.
  • Experience crafting or operationalizing multi-turn, hallucination, adversarial, or agentic-safety evaluations.

Responsibilities

  • Lead cross-functional planning and execution for evaluation identification, selection and execution. With a focus on agentic-safety evaluations, including multi-turn tool calling, unsafe tool use, unintended actions, excessive autonomy, and multi-step failures.
  • Track execution, results for all evals, analysis, and the mitigation and research plans for improvement that will stem from results analysis.
  • Translate evaluation findings into prioritized mitigation plans with accountable owners, committed dates, and measurable closure criteria
  • Lead end-to-end safety planning and execution for Nemotron models, including program scope, dependencies, risks, resources, and release-readiness criteria.
  • Translate technical safety priorities into executable programs covering LLM security, frontier risks, agentic safety, hallucinations
  • Establish and track multi-turn, multi-modal, multi-lingual, long-context, and reasoning model evaluations across relevant use cases, domains, languages, modalities, and model releases.
  • Establish governance for reviewing safety findings, assigning severity, determining release impact, and escalating unresolved risks.
  • Ensure required safety evidence, approvals, exceptions, and risk-acceptance decisions are documented before release.
  • Identify program gaps and critical dependencies early and drive decisions, escalations, recovery plans, and corrective actions.
  • Build (with AI for Code/AI for Work) and maintain dashboards and executive reporting for evaluation coverage, critical findings, mitigation progress, residual risk, and release readiness.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service