The Senior Site Reliability Engineer will be responsible for designing, implementing, and operating scalable, resilient, and highly available systems on Google Cloud Platform. This role focuses on improving service availability, latency, performance, and operational resilience. The engineer will define and track service-level indicators and objectives, perform capacity planning, and design disaster recovery capabilities. Responsibilities also include implementing secure cloud networking and IAM practices, partnering with cybersecurity teams, and optimizing cloud consumption for cost efficiency. Additionally, the role involves building and maintaining cloud infrastructure using infrastructure-as-code tools, automating operational tasks, and improving CI/CD pipelines. The engineer will develop actionable alerts, create dashboards and runbooks, participate in on-call rotations, respond to production incidents, and facilitate blameless postmortems. Collaboration with software engineering, data engineering, security, and product teams is crucial to promote shared responsibility for production reliability and establish reliability standards. The role also involves providing technical guidance on SRE, cloud, Kubernetes, observability, and incident management practices.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior