We are looking for a Senior Site Reliability Engineer with deep, hands-on experience operating highly available production environments on AWS and Amazon EKS. This is a true SRE position, not a cloud architecture, infrastructure design, or monitoring-focused role. You will take direct ownership of production reliability, participate in the on-call rotation, respond to critical incidents, troubleshoot complex Kubernetes and distributed-system failures, and drive permanent improvements following incidents. The role also requires strong technical communication. You will interact directly with customers during technical discussions and production escalations, clearly explaining issues, making sound technical decisions, and driving problems through to resolution. We are looking for someone who has spent significant time running production systems, not simply designing them.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed