Distinguished Engineer, Scaled Out Inferencing

NVIDIASanta Clara, CA
$320,000 - $488,750

About The Position

NVIDIA is leading the industry in delivering accelerated computing in cloud and enterprise environments. We’re a team of innovative engineers dedicated to solving some of the world’s biggest challenges, constantly driving advancements, and impacting millions of lives worldwide! As a technology leader at NVIDIA, you will lead the development of our global strategy for scaled-out AI inferencing. You will architect the high-throughput, low-latency distributed pipelines and model serving strategies required for massive scale and production reliability. You will define and drive the technical roadmap for full-lifecycle, from deployment and versioning to automated scaling, across enterprise and cloud environments. Working with NVIDIA leadership, you will establish the systems and orchestration layers that enable the world’s most AI models to run with peak efficiency on our accelerated computing hardware.

Requirements

  • 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration.
  • Proven track record building secure, highly available, and durable production distributed systems.
  • 7-10+ years of leadership experience
  • BS/MS or higher or equivalent experience in systems / software engineering, or related engineering fields
  • Proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.
  • Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.
  • Develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership. Ability to synthesize cross-functional needs into architecture and design while guiding internal execution across diverse teams.
  • Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high performance and accelerated computing customers.

Nice To Haves

  • Real world experience building the systems to support AI/ML workloads.
  • Direct experience in designing, developing, delivering and operating secure, highly available, scaled out systems in enterprise and cloud environments.
  • Demonstrated history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.
  • Familiarity with open source ecosystems and projects (e.g. Dynamo, TensorRT-LLM, vLLM, SGLang, Ray). Ability to collaborate and influence in open source project governance to represent NVIDIA, customers, and partners interests in technical alignment and direction.

Responsibilities

  • Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.
  • Collaborate on hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.
  • Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state of the art inferencing on NVIDIA accelerated hardware.
  • Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.
  • Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA’s solutions set the industry standard for performance and availability.
  • From ideation to architecture, design, development, deployment, operations, and full lifecycle management, lead all technical aspects of planning and continuous evolution of a large technical scope.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service