Service Reliability Engineer

NVIDIAUS, TX, Remote, TX
$168,000 - $333,500Hybrid

About The Position

At NVIDIA, our Compute Infrastructure Support (CIS) team looks for a driven Site Reliability Engineer passionate about innovation and excellence. As an SRE, you will be vital in delivering world-class support for our on-prem and cloud products and services. Your work will help maintain near 100% availability. This role offers a chance to work with powerful technology and join a diverse expert team. Together, you will build reliable systems that advance our innovative progress in AI and accelerated computing.

Requirements

  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
  • Familiarity with GPU hardware and high-performance computing environments.
  • Proficiency with observability and incident management tools (Grafana, OpenTelemetry, PagerDuty, JIRA).
  • Experience with Cloud platforms (AWS, Azure, GCP, OCI) is a plus; strong preference for on-prem expertise.
  • Highly motivated with strong communication skills, demonstrating the ability to work effectively with multi-functional teams.
  • 8+ years of experience coordinating large-scale production systems and more than 3 years of experience working in high-availability Internet, Cloud, or Data Center settings.
  • BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
  • Expert-level Linux system administration, automation using Ansible and/or Python and strong expertise in shell scripting, DNS, DHCP, storage systems, and core networking.
  • Proven experience in problem-solving and maintaining large-scale bare-metal infrastructure.
  • Excellent partnership, documentation, and mentoring skills.

Nice To Haves

  • Experience with scripting languages, particularly Python.
  • Prior experience running virtual machines under community-supported or commercial hypervisors.
  • Knowledge of application containers and container orchestration systems.
  • Basic understanding of Git.
  • Demonstrate ability to master and maintain complicated environments.

Responsibilities

  • Operate within a 24/7 follow-the-sun support model across multiple continents, working closely with a U.S.-based manager. Manage a 4-day, 10-hour schedule, including either Saturday or Sunday, with flexible early or late shifts to maintain global coverage.
  • Monitor and manage extensive production GPU and Kubernetes environments to ensure high availability and performance. Apply advanced tools to detect, prevent, and respond to incidents proactively.
  • Apply in-depth systems knowledge to analyze logs, metrics, and system behavior, diagnosing issues and implementing effective resolutions.
  • Develop and complete predictive automated support routines to prevent potential issues from impacting production.
  • Improve automation by integrating incident analysis insights into auto-healing and automated break-fix solutions.
  • Perform comprehensive systems administration, network administration, and security monitoring tasks.
  • Coordinate with domain experts and service owners to resolve complex issues efficiently.
  • Continuously improve service quality and operational processes based on incident feedback.
  • Foster strong interpersonal and communication skills to ensure effective coordination across teams during incident resolution.
  • Deliver outstanding customer focus, ensuring clients feel supported and valued throughout all interactions.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service