Staff Site Reliability Engineer

Horizon3
•$199,750 - $270,000•Remote

About The Position

Horizon3 is seeking a hands-on Staff Site Reliability Engineer to own and evolve the reliability strategy, operating model, and engineering-wide standards supporting our platform. This is a foundational role for an experienced engineer who will set technical direction across teams, lead the highest-impact reliability initiatives, and establish the practices and systems that enable engineering to operate production services safely.

Requirements

  • Experience designing, operating, and troubleshooting large scale distributed systems in production environments.
  • Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others
  • Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
  • Backend experience building backend systems and automation that reduce optional toil, strengthen safeguards, and operational efficiency.
  • Experience in leading high severity incidents and improving incident response programs.
  • Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation.
  • Python and Terraform (Infrastructure as Code), or equivalent automation and infrastructure-as-code tools.
  • Experience with Observability tools such as Datadog, New Relic, Grafana, or equivalent platforms.
  • Experience operating production services in AWS and Kubernetes
  • Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows.

Responsibilities

  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk.
  • Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness across multiple teams.
  • Establish an organization-wide approach to service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services.
  • Define and drive adoption of observability standards across pipelines and platform components that report on service health, performance and operational risk.
  • Set the standard for dashboards, actionable alerting, runbooks, and escalation paths.
  • Drive end to end complex cross-functional reliability initiatives.
  • Set and raise the engineering wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness.
  • Shape the technical direction, operationable model, and growth path of the SRE function.
  • Participate in a 24/7 on-call rotation and help design an on-call model that is sustainable, appropriately staffed, and continuously improved.

Benefits

  • health, vision & dental insurance for you and your family
  • a flexible vacation policy
  • generous parental leave
  • equity package in the form of stock options
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service