Senior Deep Learning Scientist, Multimodal Agentic RL

NVIDIASanta Clara, CA
$184,000 - $287,500

About The Position

NVIDIA is seeking Senior Deep Learning Scientists to advance its work in streaming and agentic multimodal AI. The role involves developing models capable of reasoning, planning, and acting across diverse modalities, with a focus on core algorithmic improvements for multimodal foundation models. Successful candidates will contribute to NVIDIA's Nemotron Omni and VoiceChat platforms, working on high-impact, large language models and multimodal AI products that enhance user experience. This is an opportunity to join the Nemotron LLM team and tackle real-world agentic AI challenges.

Requirements

  • Master’s degree (or equivalent experience) or PhD in Computer Science, AI, or Applied Math with 8+ years of relevant work experience.
  • Excellent programming skills in Python with strong fundamentals in scalable model development and deep learning frameworks like PyTorch.
  • Strong knowledge of ML/DL techniques and modern foundation model architectures, including Transformers and mixture-of-experts models.
  • Foundational understanding of reinforcement learning algorithms and implementation, including MDPs, policies, and reward design.
  • Hands-on experience in post-training multimodal models for omni-modality (audio-visual) reasoning, full-duplex voice chat, and human-AI interaction.
  • Proven ability to manage model development life cycles, including dataset versioning, experiment tracking, and evaluation pipelines.

Nice To Haves

  • Strong record of publications in top-tier AI and machine learning venues such as NeurIPS, ICML, ICLR, or CVPR.
  • Validated experience training and deploying multimodal foundation models using large-scale distributed infrastructure.
  • Experience applying deep reinforcement learning techniques to train multimodal agents in complex simulation or gaming environments.
  • Background in audio/speech AI, especially audio language models or audio generation.
  • Background in building embodied AI systems that integrate multimodal perception with backend action-fulfillment and long-horizon planning.

Responsibilities

  • Apply fundamental and applied research to develop, train, fine-tune, and deploy large language models for agentic systems encompassing audio-visual reasoning, tool usage, and document understanding.
  • Advance post-training and alignment methods including instruction tuning, preference optimization, and RLHF/RLVR/MOPD to improve multimodal agents for complex use cases.
  • Research and develop agentic reasoning and grounded perception capabilities, focusing on planning, tool execution, and long-horizon task completion across digital and physical environments.
  • Lead the collection, development, and benchmarking of multimodal datasets, ensuring high-quality evaluation of model accuracy, safety, and task completion success.

Benefits

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service