Research Intern, Efficient Deep Learning - 2027

NVIDIA•Santa Clara, CA
•$38 - $94

About The Position

NVIDIA is searching for an outstanding PhD intern working on efficient deep learning to join the Deep Learning Efficiency Research (DLER) team. The team has two core focuses: (1) efficient diffusion language models and multimodal generative models, and (2) efficient agentic AI with hybrid inference orchestration across cloud and edge. They are also excited about post-training model optimization (pruning, quantization, NAS), efficient architecture design, adaptive/dynamic inference, and resource-efficient training and finetuning. The intern will work within a collaborative research team that consistently publishes at top venues in computer vision and machine learning. The team's expertise includes computer vision, deep learning, generative models, diffusion LLMs, multimodal models, and hybrid cloud–edge agentic systems. Contributions have the chance to create real impact on products.

Requirements

  • Pursuing a Ph.D. in Computer Science/Engineering, Electrical Engineering, etc.
  • Excellent knowledge of theory and practice of machine learning and deep learning.
  • Experience with large language models, diffusion language models, multimodal / vision-language models, or agentic systems is required.
  • Hands-on experience with large-scale model training including data preparation and model parallelization (tensor and pipeline) is required.
  • Outstanding research track record with at least one top-tier conference (ICML, ICLR, NeurIPS, CVPR, ICCV, etc.).
  • Excellent communication skills.

Nice To Haves

  • Parallel programming (e.g., CUDA).
  • Interest or experience in hybrid cloud–edge inference, orchestration, or adaptive routing.
  • Background in pruning, quantization, NAS, or efficient backbones.

Responsibilities

  • Research, design, and implement novel methods for efficient deep learning in one or both of the team’s focus areas: Diffusion LLMs and multimodal models — sampling efficiency, adaptive unmasking, self-speculation / parallel decoding, training and distillation pipelines, and multimodal generation.
  • Research, design, and implement novel methods for efficient deep learning in one or both of the team’s focus areas: Efficient agentic AI — hybrid inference orchestration across cloud and edge, routing and scheduling policies, on-device vs. cloud expert delegation, and resource-aware agent loops.
  • Publish original research.
  • Collaborate with other team members and teams.
  • Work with product groups to transfer technology.
  • Collaborate with external researchers.

Benefits

  • Competitive salaries
  • Generous benefits package
  • Intern benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service