At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive. Role Summary We are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-glass monitoring, proactive incident response, operational maintenance, and continuous compliance support across our hybrid environment—spanning bare-metal on-premises Red Hat Enterprise Linux (RHEL) servers, AWS (Commercial and GovCloud), and Kubernetes (EKS) infrastructure operating under strict FedRAMP standards. As an SRE II, your primary responsibility is maintaining platform availability, hardware reliability, and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 hybrid platform uptime. 24/7 Shift & On-Call Expectations This position requires active participation in a 24/7/365 operational shift model: Round-the-Clock Coverage: Active "eyes-on-glass" monitoring during scheduled shifts (day, evening, night, and weekend rotations). Weekend & Holiday Rotation: Scheduled weekend shifts and holiday coverage to ensure continuous operational readiness across both on-premises data centers and cloud regions. On-Call Escalations: Primary and secondary on-call responsibilities during and outside regular shift windows to meet stringent FedRAMP Incident Response (IR) SLAs. Real-time Collaboration: Continuous presence in operational Slack channels and ChatOps bridge rooms for instant incident mobilization.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed