Senior Site Reliability Engineer

ServiceTitanUS CA Remote, CA
$137,900 - $221,400

About The Position

We're looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. This team owns the reliability and health of the applications running on top of our cloud infrastructure. They design the signals that indicate problems and build the systems that ensure ServiceTitan runs better, faster, and cheaper as we scale. The role makes a significant impact on thousands of companies globally by improving their business efficiency. The Site Reliability and Infrastructure Engineering team centralizes measurement and guidance, empowering every engineer to enhance availability and efficiency within their domain of the ServiceTitan cloud. The company fosters a culture of diversity, inclusion, and innovation, encouraging employees and their ideas to thrive.

Requirements

  • Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system.
  • SRE principles: practical experience with SLIs, SLOs, and error budgets — able to speak to how you've defined and monitored these on real systems, not just definitions.
  • Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing).
  • Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across tools.
  • CI/CD: strong understanding of a CI/CD system — GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptable.
  • Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you're most comfortable in.
  • Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures).
  • Strong production troubleshooting skills — comfortable diagnosing issues under pressure.
  • 8-10+ years of relevant hands-on experience.

Nice To Haves

  • database experience (not mandatory — databases are monitored by the same team, not owned individually).

Responsibilities

  • Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load).
  • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure.
  • Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems.
  • Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work.
  • Partner with product engineering teams to review architecture and infrastructure decisions before they ship.
  • Build and maintain automation that reduces manual, repetitive operational work across the team.
  • Write and maintain runbooks and documentation so on-call knowledge is shared across the team, not siloed with one person.
  • Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed.
  • Collaborate across engineering teams to adopt best practices in reliability and observability.
  • Contribute to CI/CD pipelines and help teams ship changes safely and quickly.

Benefits

  • Flexible time off
  • Learning and development opportunities
  • Comprehensive onboarding program
  • Leadership training
  • Programs and events
  • Bonusly
  • Peer-nominated awards
  • Company-paid medical, dental, and vision (with 100% employer paid options and 90% coverage for dependents)
  • FSA and HSA
  • 401k match
  • Telehealth options including memberships to One Medical
  • Parental leave and support
  • Up to $20k in fertility services (i.e. IUI and IVF)
  • Surrogacy
  • Adoption reimbursement
  • On demand maternity support through Maven Maternity
  • Free breast milk shipping through Maven Milk
  • Pet insurance
  • Legal advisory services
  • Financial planning tools
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service