Software Engineer Manager - Platform Reliability Engineering (Remote)

The Home DepotGEORGIA - VIRTUAL - GA01, GA
$140,000 - $240,000Remote

About The Position

As a Software Engineer Manager for Platform Reliability Engineering, you will ensure the resilience, performance, and security of our enterprise cloud foundation in line with enterprise reliability engineering standards. Working closely alongside RE and Enablement teams supporting critical Engineering Experience services, you will lead a dedicated team of engineers accountable for delivering comprehensive Reliability Engineering practices for a segment of our Cloud Platform portfolio. Your mission is to ensure that reliability is engineered into our platforms through extensive automation, rigorous change management, systematic incident and problem management, and scheduled destructive testing. You will drive operational excellence by establishing and strictly enforcing Service Level Objectives (SLOs), maintaining our security postures, and guaranteeing that product teams can build and run customer-facing workloads on highly available, paved-path solutions.

Requirements

  • Must be eighteen years of age or older.
  • Must be legally permitted to work in the United States.
  • Mastery of an object oriented programming language (preferably Java)
  • Must be legally permitted to work in the United States
  • The knowledge, skills and abilities typically acquired through the completion of a bachelor's degree program or equivalent degree in a field of study related to the job.
  • 5 years of work experience

Nice To Haves

  • 6-10 years of relevant work experience
  • Proven experience managing or leading Site Reliability Engineering (SRE), Platform Engineering, or DevOps teams.
  • Demonstrated ability to drive cultural change around process and technology adoption.
  • Experience leading incident response, driving mitigation to minimize MTTR, plus blameless post-mortems and RCAs that engineer out recurrence.
  • Experience defining and enforcing SLOs/SLIs against Critical User Journeys, using error budgets and Production Readiness Reviews to govern release velocity.
  • Strong understanding of Google Cloud or similar public cloud ecosystems.
  • Experience operating, troubleshooting, and providing tier-escalation support for complex microservice architectures and high-traffic web applications.
  • Strong understanding of modern observability stacks, container orchestration (Kubernetes/GKE), and Infrastructure as Code (Terraform).
  • Strong automation focus: chaos/resiliency testing, toil reduction via custom scripting, and automated recovery/monitoring.
  • Solid understanding of software-defined cloud networking, overlay mesh networks, and Zero Trust Network Access (ZTNA) principles.
  • Experience leading vulnerability/exposure management with automated remediation SLAs, plus peak-readiness and dynamic-scaling strategies for seasonal surges.
  • No additional education
  • No additional years of experience
  • None

Responsibilities

  • Collaborates and pairs with product team members (UX, engineering, and product management) to create secure, reliable, scalable software solutions
  • Documents, reviews and ensures that all quality and change control standards are met
  • Writes custom code or scripts to automate infrastructure, monitoring services, and test cases
  • Works with vendors and partners for the successful implementation of critical tooling and platforms
  • Creates meaningful dashboards, logging, alerting, and responses to ensure that issues are captured and addressed proactively
  • Contributes to enterprise-wide tools to drive destructive testing, automation, and engineering empowerment
  • Evaluates new technologies for adoption across the enterprise
  • Participates in and leads review board sessions to drive consistency across the enterprise
  • Fills in on product teams for engineers who are out of the office
  • Fields questions from engineers, product teams, or support teams
  • Monitors tools and participates in conversations to encourage collaboration across product teams
  • Provides application support for software running in production
  • Acts as a technical escalation point for the engineers on the team
  • Provides leadership, mentoring, and coaching to Software Engineers
  • Attracts, retains, and develops top talent to build a world class Software Engineering Team
  • Conducts annual and mid-year reviews by reviewing individual development plans and team feedback
  • Fosters collaboration with team members to drive consistency across product teams, and finds opportunities to expose engineers to career interests
  • Acts as a proponent of modern software development practices
  • Guides team members in strategy, alignment, analysis, and execution tasks within and across product teams
  • Participates in and contributes to learning activities around modern software design and development core practices (communities of practice)
  • Learns, through reading, tutorials, and videos, new technologies and best practices being used within other technology organizations
  • Builds relationships with technology leaders at other companies to learn best practices and elegant solutions to common problems

Benefits

  • The pay range for this position is between $140,000.00 - $240,000.00
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service