This role focuses on Production Support & Reliability, providing hands-on support for enterprise applications in production. The SRE will troubleshoot issues across application and infrastructure layers, drive incident triage and resolution, and ensure systems remain stable, available, and supportable. The position involves participating in a 24/7 on-call rotation, responding to and resolving production incidents, and coordinating with engineering teams during high-severity issues. Responsibilities also include executing deployments using Ansible and CI/CD pipelines, using Git workflows for release management, supporting deployment validation and troubleshooting, and improving deployment reliability. The role requires monitoring systems using tools like Splunk, Google Analytics, CloudWatch, and synthetic monitoring tools, analyzing logs, metrics, and alerts, and contributing to dashboards. Additionally, the SRE will manage installation, patching, and upgrades of third-party applications, maintain systems aligned with N-1 patching standards, and perform ongoing maintenance. Automation is a key aspect, involving the use and maintenance of existing Ansible automation and scripts, and improving operational processes through targeted automation. Collaboration with Cloud Engineering, Development, and Security teams is essential, as is supporting the onboarding of new applications and maintaining documentation and runbooks.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed