About The Position

As a CaaS Private Site Reliability Engineer, you will lead reliability, resilience, and operational excellence for the CaaS Private platform in US. You will bring production engineering discipline to Kubernetes, observability, automation, and incident management while helping teams improve service health and platform readiness. You will partner across engineering, operations, and application teams to strengthen SLOs, reduce manual intervention, and ensure platform changes are measurable, supportable, and aligned to enterprise reliability standards.

Requirements

  • Extensive experience with Kubernetes, Linux, distributed systems reliability, and production platform operations
  • Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root cause analysis
  • Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements
  • Experience defining SLOs, service indicators, alert quality standards, production readiness practices, and escalation models
  • Ability to lead complex reliability improvements independently while partnering across engineering, operations, and application teams

Nice To Haves

  • Strong communication skills with the ability to explain technical findings to technical and non-technical stakeholders
  • Sound operational judgment with the ability to balance urgency, risk, and long-term platform stability during incidents
  • Growth mindset with a focus on continuous learning, process improvement, and measurable reliability outcomes
  • Experience mentoring engineers and improving operational culture across platform or infrastructure teams
  • Background with automation tools, infrastructure-as-code practices, capacity planning, disaster readiness, or cloud-native platform operations

Responsibilities

  • Lead the reliability strategy for the CaaS Private platform in US, including SLO frameworks, operational standards, and incident management maturity
  • Drive resilience improvements across observability, capacity planning, upgrade safety, disaster readiness, and operational automation
  • Own complex production management issues by leading troubleshooting, identifying root causes, and implementing preventive fixes
  • Define and refine service indicators, alert thresholds, dashboard standards, production readiness criteria, escalation paths, and postmortem follow-through
  • Develop automation and self-healing workflows that reduce manual intervention, improve recovery times, and strengthen platform supportability
  • Partner with cross-functional teams to ensure platform changes are measurable, supportable, and aligned with reliability objectives
  • Mentor junior and middle engineers while fostering a culture of blameless learning, measurable reliability, and operational excellence
  • Influence platform architecture and roadmap decisions using data from incidents, capacity models, operational trends, and reliability metrics
  • Communicate strategic insights and practical recommendations to engineering, operations, and business stakeholders to guide reliability investments

Benefits

  • A hybrid working model, allowing for in-office / work from home flexibility
  • Generous vacation, personal and volunteer days
  • Employee Resource Groups support an inclusive workplace for everyone and promote community engagement
  • Competitive compensation packages
  • Health and wellbeing benefits
  • Retirement savings plans
  • Parental leave
  • Family building benefits
  • Educational resources
  • Matching gift and volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service