Luma is seeking a Senior Site Reliability Engineer to own the GPU infrastructure that powers its research and product. This role involves managing thousands of NVIDIA and AMD GPUs across on-premise and multi-cloud environments (AWS and OCI). The Senior SRE will be responsible for ensuring the reliability and speed of training and inference clusters, and will contribute to redesigning these systems for future scalability. This is a hands-on, low-level role for a Linux engineer who can handle complex GPU, networking, and kernel-level failures, including direct collaboration with vendors like NVIDIA. The position is ideal for someone who thrives on solving intricate problems in a dynamic, less-structured environment, and is not suited for those seeking a narrowly defined operational role.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed