Senior Site Reliability Engineer

RADAR•San Diego, CA
•$170,000 - $219,000•Remote

About The Position

RADAR runs data infrastructure across 1,600+ live retail stores, processing tens of billions of real-world events every day. We’re hiring a Site Reliability Engineer to own the reliability of that system end to end — leading incident response, running day-to-day NOC operations, and building the observability foundation that lets us catch issues before they hit a store floor. You’ll be the steady hand during a live incident, and the engineer making sure there are fewer of them to begin with.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or Production Operations, with direct incident response and on-call experience.
  • Experience running or actively contributing to a NOC, including shift scheduling, escalation processes, and performance metrics.
  • Strong hands-on experience with observability tooling (Prometheus, Grafana, Datadog, New Relic, Splunk, ELK, OpenTelemetry, or similar).
  • Solid understanding of SLIs, SLOs, SLAs, and error budgets, and how to use them to drive prioritization.
  • Hands-on release engineering experience, including CI/CD pipelines, deployment automation, and safe rollout practices like canary releases, feature flags, and automated rollbacks.
  • Proficient in at least one scripting or programming language (Python, Go, Bash, etc.).
  • Experience with infrastructure-as-code tools (Terraform, Ansible).
  • Experience with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker).
  • Clear, direct communicator who stays calm and organized under pressure during live incidents.

Nice To Haves

  • Experience building or scaling a NOC from the ground up.
  • Background in distributed systems architecture and microservices troubleshooting.
  • Familiarity with chaos engineering and resilience testing.
  • Certification such as AWS Certified SysOps Administrator, Google Professional Cloud DevOps Engineer, or ITIL.

Responsibilities

  • Own the incident management lifecycle end to end: detection, triage, escalation, communication, resolution, and postmortem for production incidents.
  • Act as Incident Commander for high severity incidents, coordinating across engineering, support, and leadership until resolution.
  • Run day-to-day NOC (Network Operations Center) operations, including 24/7 shift coverage, escalation matrices, and shift handover protocols.
  • Coach and mentor NOC analysts on triage discipline and escalation judgment, and own NOC KPIs like response time and escalation accuracy.
  • Design and maintain observability pipelines across metrics, logs, and traces, and define SLIs/SLOs with engineering and product.
  • Build dashboards and alert that surface true signal from our sensor and platform data, cutting down on noise and alert fatigue.
  • Facilitate blameless postmortems and root cause analysis, and track corrective actions through to closure.
  • Maintain on-call rotations, runbooks, and escalation policies, and report on MTTA/MTTR/MTBF trends to leadership.

Benefits

  • equity
  • comprehensive medical and dental coverage
  • life and disability benefits
  • 401k plan
  • flexible time off
  • paid parental leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service