About The Position

Running GPU-accelerated Kubernetes reliably is deceptively hard. A small change to a driver, kernel, operator, or Kubernetes version can break a cluster in ways that are painful to diagnose and expensive to reproduce. AI Cluster Runtime (AICR) is NVIDIA's open source project to define a validated, reproducible software stack that can go into every AI cluster, published as version-locked recipes customers can trust through end-to-end supply-chain provenance. NVIDIA is hiring a product manager to lead the accelerated runtime that defines every GPU Kubernetes cluster. You'll set the direction for the NVIDIA AICR project and grow it into how the industry runs NVIDIA's software stack on Kubernetes, growing its coverage across accelerators, clouds, and workloads. The role also represents AICR to customers, partners, and the open-source community. Want to bring GPU infrastructure to life and power NVIDIA innovation? We'd love to hear from you!

Requirements

  • 12+ years in technical product management across cloud infrastructure, Kubernetes, or developer and platform tooling.
  • A bachelor's or master's degree in a relevant field, or equivalent experience.
  • Depth in the Kubernetes GPU stack: cluster provisioning and lifecycle, operators and drivers, scheduling, and the compatibility failure modes across kernel, driver, OS, and Kubernetes versions.
  • Experience with GitOps and cluster configuration delivery (Helm, Argo CD, Flux, or equivalent tools) and the realities of shipping configuration that has to be reproducible.
  • A track record of contributing to open-source projects, including community strategy and external presence such as conference talks or technical writing.

Nice To Haves

  • Hands-on work with NVIDIA GPU infrastructure, DGX systems, the GPU Operator, or GPU-aware Kubernetes scheduling.
  • Software supply-chain security: SLSA, SBOMs, Sigstore or Cosign, or reproducible builds.
  • Packaging and distributing a sophisticated software stack (operators, drivers, platform components) across clouds and hardware.
  • Owning or contributing to a widely adopted cloud-native or CNCF open-source project.

Responsibilities

  • Own AICR and related technologies, set product direction, and standardize how customers and partners run NVIDIA's GPU software stack on Kubernetes.
  • Expand coverage across accelerators, clouds, and workloads, from training and inference to agentic AI.
  • Keep the stack current and trustworthy, turning NVIDIA's validation into recipes that customers can deploy with supply-chain confidence.
  • Represent AICR and other NVIDIA projects to customers and partners, and lead their growth in the open-source community.

Benefits

  • highly competitive salaries
  • comprehensive benefits package
  • equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service