NVIDIA is a leader in AI, High Performance Computing, and Visualization. The company is seeking a motivated Deep Learning engineer to integrate advanced CUDA features and Distributed Runtime technologies into AI stacks like PyTorch, TRT-LLM, vLLM, SGLang, and JAX. The role involves working with the team that developed core CUDA features for scaling Deep Learning and HPC applications. The engineer will address diverse multi-GPU demands, from large-scale training to microsecond latency inference, aiming to improve both productivity and performance of AI applications. This is an opportunity for individuals with an AI background to advance the state of the art in the field and contribute to NVIDIA's vision.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Ph.D. or professional degree