As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems. You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives. Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management. You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal
Education Level
No Education Listed