We are seeking a Senior Site Reliability Specialist to join our SRE team. Reporting to the Engineering Manager, you are responsible for scaling AWS cloud infrastructure, evolving Kubernetes deployment pipelines, improving monitoring, alerting, and resiliency, and developing tooling that enables product teams to deliver safely and efficiently. For acquired Azure-based products the focus is on monitoring, alert triage, and runbook-driven incident response rather than greenfield platform design. This role owns shared platform services across cloud regions, including databases, messaging, logging, search, and tenant provisioning. The Senior SRE is expected to lead major infrastructure initiatives and proof-of-concept efforts, contribute to technical planning and prioritization, as well as partners with the Product teams to reduce operational incidents. The SRE team also develops and operates AI-driven tools to streamline runbooks, accelerate incident response, and generate operational insights from platform telemetry to improve reliability and reduce manual effort.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior