The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites. The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience. It owns production reliability and resilience, observability architecture, service-level objectives, incident response, and implementation of AIOps workflows for triage, remediation, and self-healing. As a Senior Staff Service Reliability and Operational Intelligence Engineer, you define the technical direction for reliability across regions and services. You own the reliability strategy, establish the standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. You stay deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation. The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed