As a Senior Site Reliability Engineer (SRE), you will lead the design and management of large-scale, highly available distributed systems, collaborating with application teams to ensure the performance, reliability, and security of our SaaS platform. You will deploy resilient AWS cloud-native services and leverage CNCF-standard tools—such as Kubernetes, Prometheus, and Service Mesh—to standardize operations across our multi-region, microservice-based architecture. A core component of your work involves developing automation for service operations, including deployment, chaos testing, and "everything-as-code" strategies, to provide robust guardrails for our rapidly growing infrastructure. You will directly influence the ThousandEyes platform by identifying and resolving operational obstacles, ensuring our systems remain scalable under substantial daily data volumes. This role is exceptionally exciting because it places you at the center of mission-critical engineering, where you will solve complex scaling challenges and shape the future of our global platform's reliability.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior