Engineer, Site Reliability

T-Mobile USAAtlanta, GA
Onsite

About The Position

This role is essential for maintaining and improving the reliability and resilience of digital infrastructure systems. It primarily involves automating processes, monitoring system health, and managing incident responses to reduce operational disruptions. The role requires proficiency in programming, scripting, and incident management to support system stability and efficiency. Success is measured by system uptime, reduction in manual interventions, and rapid recovery from incidents. The work directly supports organizational service quality and operational performance by ensuring robust and reliable digital operations.

Requirements

  • Bachelor's Degree plus 3 years of related work experience OR advanced degree with 1 year of related work experience OR combination of education and experience deemed equivalent (Required)
  • Acceptable areas of study include Computer Science or Engineering
  • Application Monitoring (Required)
  • Automation (Required)
  • CI/CD (Required)
  • Capacity Planning (Required)
  • Cloud Computing (Required)
  • Incident Management (Required)
  • Performance Tuning (Required)
  • Scripting (Required)
  • System Reliability (Required)
  • At least 18 years of age
  • Legally authorized to work in the United States

Nice To Haves

  • Master's/Advanced Degree Computer Science or Data Science (Preferred)
  • 2-4 years Developing and maintaining CI/CD pipelines for software deployment (Preferred)
  • 2-4 years Implementing and managing cloud-native platforms and solutions (Preferred)
  • 2-4 years Guiding and mentoring teams in reliability engineering practices (Preferred)
  • Certified Kubernetes Administrator (CKA) - Certification that validates the ability to use Kubernetes, which is crucial for automating deployment, scaling, and operations of application containers across clusters of hosts. (Preferred)
  • AWS Certified DevOps Engineer - Certification that demonstrates an individual's expertise in provisioning, operating, and managing distributed application systems on the AWS platform. (Preferred)
  • Site Reliability Engineering (SRE) Foundation Certification - Certification that provides a foundational understanding of the SRE philosophy, practices, and tools to enhance the reliability and performance of systems. (Preferred)

Responsibilities

  • Automate processes to improve system reliability and reduce manual operational tasks
  • Monitor systems proactively to minimize operational incidents and maintain service continuity
  • Streamline software development and deployment processes to enhance operational efficiency
  • Develop scripts and tools to decrease manual efforts in routine operational activities
  • Manage incident response to ensure rapid recovery and minimize service disruption
  • Adapt to new technologies to sustain and improve system robustness and performance
  • Also responsible for other duties/projects as assigned by business management as needed

Benefits

  • Competitive base salary and compensation package
  • Annual stock grant
  • Employee stock purchase plan
  • 401(k)
  • Access to free, year-round money coaches
  • Medical, dental and vision insurance
  • Flexible spending account
  • Paid time off
  • Up to 12 paid holidays
  • Paid parental and family leave
  • Family building benefits
  • Back-up care
  • Enhanced family support
  • Childcare subsidy
  • Tuition assistance
  • College coaching
  • Short- and long-term disability
  • Voluntary AD&D coverage
  • Voluntary accident coverage
  • Voluntary life insurance
  • Voluntary disability insurance
  • Voluntary long-term care insurance
  • Mobile service & home internet discounts
  • Pet insurance
  • Access to commuter and transit programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service