About The Position

NVIDIA researchers depend on GPU clusters for large-scale AI workloads. Our DGX Cloud Kubernetes Runtime & Release team brings those clusters to life across major public clouds and specialized GPU providers, often on hardware that is new to the world when we get it. We build and maintain the supported Kubernetes runtime, automate its delivery, and bring new providers and GPU platforms into production. We’re growing quickly and taking on broader ownership of NVIDIA’s cluster software delivery. We’re hiring across Runtime, Release Engineering, and Provider Integration, with each role focused on your strengths. You don’t need experience across every area below.

Requirements

  • 6+ years building production infrastructure software or distributed systems.
  • Strong programming skills in Go or another language to build production systems, with willingness to work primarily in Go.
  • Kubernetes experience and depth in at least one area: controllers and operators, release automation, test and validation systems, or cloud integration.
  • Experience delivering engineering projects, diagnosing complex failures, and collaborating across teams.
  • BS or MS in Computer Science, Engineering, or equivalent experience.

Nice To Haves

  • Go development with controller-runtime, CRDs, and reconcilers.
  • Release qualification across multiple environments or platforms.
  • GPU infrastructure, accelerated networking, or GPU scheduling.
  • Bringing new hardware, regions, or cloud providers into production.
  • Resource allocation, leasing, or fair-share scheduling and upstream integration, compatibility, or software supply chain integrity.

Responsibilities

  • Build Go controllers and APIs to install, upgrade, and validate GPU cluster software.
  • Integrate components, define API contracts, and evolve Helm and Argo CD delivery toward controller-driven automation.
  • Build validation pipelines that inform release decisions across providers and GPU platforms.
  • Develop systems to allocate GPU capacity across validation runs and account for cloud reservations and quotas.
  • Make qualification more efficient through reusable tests and clear failure reports.
  • Bring new providers and GPU hardware into production, potentially among the first engineers working with new silicon.
  • Resolve integration failures with partner teams and turn initial provisioning, upgrade, and operational checks into repeatable automation.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service