NVIDIA is searching for an outstanding PhD intern working on efficient deep learning to join the Deep Learning Efficiency Research (DLER) team. The team has two core focuses: (1) efficient diffusion language models and multimodal generative models, and (2) efficient agentic AI with hybrid inference orchestration across cloud and edge. They are also excited about post-training model optimization (pruning, quantization, NAS), efficient architecture design, adaptive/dynamic inference, and resource-efficient training and finetuning. The intern will work within a collaborative research team that consistently publishes at top venues in computer vision and machine learning. The team's expertise includes computer vision, deep learning, generative models, diffusion LLMs, multimodal models, and hybrid cloud–edge agentic systems. Contributions have the chance to create real impact on products.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Intern
Education Level
Ph.D. or professional degree