This role focuses on ensuring the reliability, availability, scalability, and performance of large-scale distributed systems. The Site Reliability Engineer will work closely with Infrastructure and Development teams to maintain the ADT platform and protect customers. This involves collaborating with cross-functional partners to improve operational health and implement SRE best practices. The position requires driving operational excellence through problem-solving, performance enhancements, and building resilient production environments. Key responsibilities include using tools like Terraform, Ansible, Kubernetes, and Dynatrace, working within cloud environments (AWS, GCP), identifying and resolving performance bottlenecks, building and maintaining infrastructure as code, contributing to observability and monitoring, supporting software releases, and providing production support, including on-call participation.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed