Sr. SRE

Pura•Pleasant Grove, UT
•Hybrid

About The Position

As a senior SRE, you'll be a hands-on driver of infrastructure reliability and delivery velocity for our Web, Mobile, Backend, and Data teams—automating away toil, hardening our systems, and mentoring engineers along the way. We're hiring with flexibility across two tracks. Depending on your background and our team's needs, you'll focus primarily on one: Track A: Infrastructure & DevOps, or Track B: Observability.

Requirements

  • Hands-on Kubernetes administration (Helm, networking, troubleshooting) and CI/CD tooling (GitHub Actions, ArgoCD, or similar) required.
  • 5–8 years as an SRE, DevOps, or Observability Engineer supporting production systems at scale.
  • Proficiency in Python, Go, or Node.js; solid Linux/bash fundamentals.
  • Experience with AWS and/or GCP and working knowledge of Kubernetes.
  • Strong troubleshooting skills and a calm, methodical approach to production issues.
  • Clear communicator who partners well across engineering teams.
  • Hands-on experience with LGTM, Prometheus, Datadog, New Relic, or ELK required.

Nice To Haves

  • Compliance support experience (SOC 2, PCI) is a plus.
  • optionally extend to our device/IoT fleet.

Responsibilities

  • Build and operate our Kubernetes platform (upgrades, autoscaling, networking, workload tuning) and CI/CD pipelines that get code to production fast and safely.
  • Own Infrastructure as Code (Terraform) across AWS and/or GCP, including IAM design and least-privilege access patterns.
  • Manage secrets and identity: AWS Secrets Manager, OpenBao, Kubernetes Secrets, External Secrets Operator, SPIFFE/SPIRE, OIDC federation.
  • Support stateful services in production, enforce guardrails via Policy-as-Code, and manage DNS/edge/CDN configuration.
  • Improve deployment safety (progressive delivery, automated rollbacks) and drive incident root-cause analysis and remediation.
  • Build and maintain observability pipelines on the LGTM stack (Loki, Grafana, Tempo, Mimir/Prometheus), instrumented via OpenTelemetry.
  • Own telemetry cost and cardinality management as our systems scale, and manage observability infra as code (Terraform, Helm).
  • Partner with Mobile and Web to build client-side/user-facing telemetry; optionally extend to our device/IoT fleet.
  • Establish log hygiene (including PII handling), and own alerting/paging tooling (PagerDuty, Opsgenie, or similar).
  • Build dashboards, SLOs, and alerting strategies that surface signal and support cost-aware infrastructure decisions; support capacity planning and cloud cost/FinOps monitoring.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service