About The Position

The Site Reliability Engineering Team Lead (Principal SRE) leads Cerence's Site Reliability Engineering team, owning the reliability, availability, and operational health of our cloud-native automotive AI platform. This role combines technical leadership of the team and the function with deep technical credibility: you will help select and mentor the team, define the reliability roadmap, govern SLI/SLO/SLA targets, and champion a blameless culture. You will serve as the Tier 2 technical escalation point for major incidents, partner with development leadership to embed reliability into the SDLC, and set strategic direction for observability and automation. This is a Principal-level role on our technical track, with no direct reports: you lead the team and own the function through technical authority. The ideal candidate brings a strong hands-on background in SRE, cloud platforms, and container orchestration, and a track record of leading site reliability teams.

Requirements

  • 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function
  • A track record of setting technical direction and holding standards across a team — with or without formal authority
  • Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio)
  • Experience with public cloud platforms (Azure primarily; AWS and Google Cloud)
  • Familiarity with observability tooling — metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana)
  • Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux)
  • Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.)
  • Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS)
  • Excellent written and verbal communication skills in English

Nice To Haves

  • Previous site reliability leadership experience
  • Experience leading distributed or multi-site technical teams
  • Background in high-availability service design (redundancy, failover, blast radius)
  • Experience with log aggregation and analytics platforms (Loki, Thanos)
  • Familiarity with ITSM and project tooling (Jira, Confluence)
  • Experience in automotive, embedded, or latency-sensitive production environments

Responsibilities

  • Help select, mentor, and technically develop the team, across multiple locations
  • Set technical direction and priorities for the team, and contribute performance and growth input to their managers
  • Design and maintain a sustainable on-call rotation; monitor page load and team health as first-class concerns
  • Own and drive the team's reliability roadmap across a 2–3 quarter horizon
  • Define and govern SLI/SLO/SLA frameworks to hold the contracted availability targets for our customer programs, which run as high as 99.95%
  • Serve as the Tier 2 technical escalation point for major incidents, partnering with the Global Operations Center, which owns incident management and response
  • Champion blameless postmortem culture — model it, reinforce it, and ensure it produces actionable outcomes
  • Lead and continuously improve Production Readiness / NFR reviews with development teams
  • Contribute to root cause analysis and own the systemic improvements that come out of it
  • Act as a named approver for high-risk and out-of-window production change
  • Set the strategic direction for metrics, dashboards, and alerting across SLI/SLO, escalation, and automation layers
  • Drive development and adoption of CI/CD automation pipelines for service deployments, rollbacks, and operational tasks
  • Partner with DevOps and platform teams to evolve shared infrastructure
  • Partner with development managers and architects to embed reliability into the SDLC by default
  • Participate in service reliability consulting and architectural reviews
  • Communicate reliability posture and risk clearly to technical and non-technical stakeholders

Benefits

  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service