Site Reliability Engineer

Runpod
$150,000 - $200,000Remote

About The Position

Runpod is the AI Developer Cloud, serving over one million developers who use the platform to experiment, train, fine-tune, deploy, and scale AI. The platform has processed more than 20 billion inference requests. Runpod is at an inflection point for AI infrastructure and is building the platform for the next generation of developers. The company is a small, remote-first team that values ownership, speed, and shipping impactful work. The Reliability team is responsible for the availability, performance, and operational excellence of Runpod’s global platform, ensuring systems are resilient, observable, and scalable. This role blends software engineering with production operations, focusing on reliability frameworks, SLO design, automation, and production hardening to reduce errors and improve performance across services and infrastructure. It's a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.

Requirements

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background check

Nice To Haves

  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks

Responsibilities

  • Define and implement SLIs/SLOs for critical services
  • Lead incident response and coordinate cross-team mitigation efforts
  • Conduct blameless postmortems and ensure corrective actions are completed
  • Perform production readiness reviews for new services and features
  • Identify systemic risks and drive preventative improvements
  • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
  • Improve signal-to-noise ratio in alerts and reduce alert fatigue
  • Build internal tooling for reliability tracking and reporting
  • Improve visibility into GPU performance and distributed systems health
  • Automate recurring operational workflows
  • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
  • Improve deployment safety through automation and guardrails
  • Strengthen CI/CD reliability and release processes
  • Partner with engineering teams to improve system resilience
  • Provide guidance on fault tolerance, scalability, and failure handling
  • Contribute to architectural discussions with a reliability-first mindset

Benefits

  • Competitive base pay ($150,000-$200,000 USD)
  • Meaningful equity (stock options)
  • Generous medical, dental & vision plans
  • Flexible PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service