SpaceX HPC is a shared compute platform used across the company for vehicle and structures simulation, machine learning, AI inference, and more. This role supports every program at SpaceX to design and operate the world's most advanced rockets and satellites. The position focuses on implementing a Site Reliability Engineer operating model for these capabilities, aiming to reduce toil, increase automation, improve observability, and establish a sustainable incident process to accelerate world-class engineering at SpaceX. The ideal candidate will own the entire HPC ecosystem, from Linux machines and Infrastructure as Code to storage and user-facing applications, treating it as a product rather than a ticket queue. While prior HPC experience is not required, strong production instincts, the ability to write code to eliminate toil, and a focus on user productivity are essential. The role involves collaborating with HPC systems engineers who design and commission clusters to enhance reliability and provide world-class services for engineers.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level