As a(n) SRE/ Site Reliability Engineer you will: Define and govern SLIs, SLOs, and SLAs for supported services. Centralized ownership of reliability practices across multiple applications and platforms. Lead incident management, major incident response, postmortems, and root cause analysis. Implement and maintain monitoring, observability, and alerting solutions using enterprise-standard tools. Drive toil reduction and automation through scripting and engineering solutions. Manage and improve AWS/cloud infrastructure reliability. Build and maintain CI/CD pipelines and infrastructure automation using tools such as Terraform. Perform capacity planning, performance tuning, and resiliency testing. Partner with application teams to improve reliability, availability, scalability, and operational excellence. Establish SRE best practices based on principles from the Google SRE Handbook. Act as an internal consulting team, helping development teams adopt reliability standards rather than owning application feature development.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed