We are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-glass monitoring, proactive incident response, operational maintenance, and continuous compliance support across our hybrid environment—spanning bare-metal on-premises Red Hat Enterprise Linux (RHEL) servers, AWS (Commercial and GovCloud), and Kubernetes (EKS) infrastructure operating under strict FedRAMP standards. As an SRE II, your primary responsibility is maintaining platform availability, hardware reliability, and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 hybrid platform uptime.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed