Senior Site Reliability Engineer

LumaRedwood City, CA
Remote

About The Position

Luma is seeking a Senior Site Reliability Engineer to own the GPU infrastructure that powers its research and product. This role involves managing thousands of NVIDIA and AMD GPUs across on-premise and multi-cloud environments (AWS and OCI). The Senior SRE will be responsible for ensuring the reliability and speed of training and inference clusters, and will contribute to redesigning these systems for future scalability. This is a hands-on, low-level role for a Linux engineer who can handle complex GPU, networking, and kernel-level failures, including direct collaboration with vendors like NVIDIA. The position is ideal for someone who thrives on solving intricate problems in a dynamic, less-structured environment, and is not suited for those seeking a narrowly defined operational role.

Requirements

  • 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
  • Working experience with Terraform, Airflow, and Ray.
  • Strong experience with AWS or OCI.
  • Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
  • Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
  • Comfort in a less-structured, fast-paced environment.

Nice To Haves

  • Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
  • Experience managing large-scale GPU clusters for AI/ML training or inference.
  • Familiarity with Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in data pipelines and infrastructure.

Responsibilities

  • Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
  • Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply, at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
  • Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
  • Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service