About The Position

NVIDIA Base Command Manager powers thousands of clusters worldwide, varying from a few to several thousands of nodes, and streamlines cluster provisioning, workload management, and infrastructure monitoring. It provides all the tools you need to deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters with NVIDIA solutions AND internal clusters used for research, operations, and next-generation projects.

Requirements

  • Bachelor's Degree or equivalent experience in Computer Science or related field.
  • 8+ years of experience in site reliability engineering and/or software development roles.
  • Fluency in Python
  • In-depth knowledge of Linux and networking

Nice To Haves

  • Experience with C++, high-performance computing, Kubernetes and/or system administration would be an asset
  • Previous experience as a system admin running BCM/Bright Cluster Manager/Base Command Manager clusters is a definite plus.
  • Proficiency with cluster networking including InfiniBand and Spectrum-X

Responsibilities

  • Contributing to deployments and daily operations of large scale next-generation GPU platforms
  • Handling incidents in GPU clusters, bridging the gap between cluster operations and development
  • Designing and implementing small features in the Base Command Manager product to become intimately familiar with the workings of the product
  • Validating complex cluster configurations including Slurm and Kubernetes orchestrators for performance, scalability and resilience, ensuring they meet real-world customer scenarios.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service