The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the platform stable and reliable, with a strong focus on service continuity and customer experience. It owns production reliability, service-level objectives, observability architecture, backup and disaster recovery, incident response, and resilience, and co-owns cloud security posture and runtime vulnerability management with DevSecOps. As Staff Site Reliability Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing. The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed