Site Reliability Engineer Jobs

1,132 jobs found — updated daily

Site Reliability Engineer

Incident IQAlpharetta, GA
Onsite

About The Position

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what “reliable” means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships. Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people. We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking.
  • Actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it.
  • A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role.
  • Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
  • Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
  • Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
  • Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
  • Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
  • Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something “can't” be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
  • You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.

Nice To Haves

  • Experience standing up an SRE practice from zero to one (“founding SRE”).
  • Experience with GitOps and just-in-time production access models.
  • Familiarity with eBPF-based auto-instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eBPF Instrumentation, for legacy or hard-to-modify codebases.
  • .NET experience is a plus, given our engineering stack.
  • Certified Kubernetes Administrator (CKA) or Google Cloud Professional DevOps Engineer.

Responsibilities

  • Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

Benefits

  • medical
  • dental
  • vision
  • life insurance
  • 401k match
  • paid-time off (PTO)

Career Resources

Build a Resume for Site Reliability Engineer

The resume builder that gets results.

  • Get clear feedback so you look as qualified as you are
  • Align your resume with the job to get further in the process, faster
  • Take the guesswork out of resume writing

Explore Related Job Searches

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service