Senior Site Reliability Engineer

InspireAtlanta, GA
Hybrid

About The Position

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These roles blend software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale. The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.

Requirements

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Science or related field
  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

Nice To Haves

  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms

Responsibilities

  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact
  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation
  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)
  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance
  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service