Reliability Engineering ● Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. ● Improve service availability, latency, performance, scalability, and operational resilience. ● Define, implement, and track service-level indicators, service-level objectives, and error budgets. ● Perform capacity planning, performance analysis, and workload forecasting. ● Design and validate disaster recovery, backup, failover, and service-restoration capabilities. ● Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices. ● Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards. ● Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency. ● Identify operational risks and recommend improvements to cloud architecture and service design. Automation and Platform Engineering ● Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools. ● Automate repetitive operational activities and systematically identify, measure, and reduce manual toil. ● Build reusable infrastructure modules, deployment patterns, and operational tooling. ● Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery. Observability and Incident Management ● Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise. ● Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation. ● Participate in a sustainable on-call rotation supporting production systems. ● Respond to production incidents, coordinate service restoration, and lead incident response when appropriate. ● Facilitate blameless postmortems and identify corrective and preventive actions. ● Use incident and operational data to improve system design, automation, monitoring, and response processes. Collaboration and Service Ownership ● Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations. ● Promote shared responsibility for production reliability between application development and platform teams. ● Establish and document reliability standards, operational practices, and reusable engineering patterns. ● Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior