About The Position

NVIDIA is seeking a motivated system software engineer with a deep understanding of device drivers, memory coherency & consistency models, strong C/C++ skills, and an interest in multi-node scalability. This role involves working on the CUDA Driver, a core component of NVIDIA's platform for accelerating general-purpose computation on the GPU. The team delivers features and improvements to realize the potential of NVIDIA hardware for various computational workloads, including deep learning, scientific computation, data science, self-driving cars, video games, and virtual reality. As a member of the team, you will use your design abilities, coding expertise, and creativity to deliver the best compute platform in the world, crafting elegant solutions to exciting problems and shaping the future direction of CUDA through collaboration with peers across NVIDIA.

Requirements

  • BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience)
  • Strong C and C++ programming skills
  • Minimum of 8 years of related development experience
  • Experience driving projects across multiple teams
  • Experience working with large codebases
  • Background with operating system interfaces for threads, process control, and virtual memory
  • Experience writing and debugging multithreaded programs
  • Good written communication as well as presentation skills

Nice To Haves

  • Prior experience with parallel computing, PyTorch, low-latency AI inference
  • Understanding of system level architecture, such as interconnects, memory hierarchy, interrupts, and memory-mapped IO
  • Knowledge of memory coherence and consistency models
  • Background with kernel mode development
  • Experience with Linux, or Windows Systems Software development

Responsibilities

  • Evangelize, architect, and implement new features related to CUDA’s memory model and multi-node scalability geared towards next-gen AI applications and deployments
  • Coordinate and drive development efforts across multiple teams
  • Help define forward-looking improvements to the CUDA APIs and programming model
  • Write effective, maintainable, and well-tested code
  • Develop code for multiple operating systems

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service