Site Reliability Engineer (High Performance Computing)

SpaceXHawthorne, CA
$125,000 - $195,000Onsite

About The Position

SpaceX HPC is a shared compute platform used across the company for vehicle and structures simulation, machine learning, AI inference, and more. This role supports every program at SpaceX to design and operate the world's most advanced rockets and satellites. The position focuses on implementing a Site Reliability Engineer operating model for these capabilities, aiming to reduce toil, increase automation, improve observability, and establish a sustainable incident process to accelerate world-class engineering at SpaceX. The ideal candidate will own the entire HPC ecosystem, from Linux machines and Infrastructure as Code to storage and user-facing applications, treating it as a product rather than a ticket queue. While prior HPC experience is not required, strong production instincts, the ability to write code to eliminate toil, and a focus on user productivity are essential. The role involves collaborating with HPC systems engineers who design and commission clusters to enhance reliability and provide world-class services for engineers.

Requirements

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree
  • 2+ years of experience with Linux operating systems in production
  • 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own

Nice To Haves

  • 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering
  • Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar)
  • Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar)
  • Experience writing scripts/code (eg. Python or similar languages) to automate common tasks
  • Experience with containers (Docker, Podman, Singularity/Apptainer)
  • Experience with Kubernetes administration for on-premise deployment
  • Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management
  • Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute — not required; we will teach this
  • Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads
  • Good understanding of version control, testing, continuous integration, build, deployment and monitoring
  • Ability to communicate clearly with users, peers, and vendors in both incident and design settings
  • Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
  • Eligibility for access to classified material up to TS/SCI with polygraph

Responsibilities

  • Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack
  • Build observability for both HPC administrators and end users — cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platform
  • Reduce toil with automation; split time between operating production systems and writing the software that makes that work smaller
  • Sustainably manage resources, including compute and storage
  • Lead capacity planning with users across the company: understand what they will need next, and turn that into a concrete picture of tomorrow's compute and storage
  • Collaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructure

Benefits

  • Long-term incentives, in the form of company stock or long-term cash awards
  • Potential discretionary bonuses
  • Employee Stock Purchase Plan
  • Comprehensive medical, vision, and dental coverage
  • 401(k) retirement plan
  • Short and long-term disability insurance
  • Life insurance
  • Paid parental leave
  • Various other discounts and perks
  • 3 weeks of paid vacation
  • 10 or more paid holidays per year
  • Paid sick leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service