Staff Software Engineer - Customer Reliability Engineering (Remote)

The Home DepotGEORGIA - VIRTUAL - GA01, GA
$90,000 - $190,000Remote

About The Position

The Customer Reliability Engineering team ensures the continuous performance, security, and resilience of various customer profile and communication services that support the company's e-commerce platform and in-store systems. The Staff Reliability Engineer is responsible for leading the design and implementation of foundational systems that ensure the reliability, scalability, performance, and efficiency of the products our customers and associates love. As a Staff SRE, you will serve as a technical anchor, establishing the blueprints for observability, automation, and cloud infrastructure architecture across the enterprise. You will drive tool selection, configuration, security, resilience, performance tuning, and production monitoring to systematically reduce operational toil. In this role, you will balance rapid product delivery with long-term system stability by defining Service Level Objectives (SLOs), error budgets, and driving a culture of blameless post-incident reviews. As a core player on the engineering team, you are expected to drive technical decisions across teams without formal authority and actively mentor engineers of all experience levels to elevate the organization's engineering standards.

Requirements

  • Must be eighteen years of age or older.
  • Must be legally permitted to work in the United States.
  • 8+ years of relevant professional experience in Cloud Operations, Site Reliability Engineering, DevOps, or Software Engineering in a high-scale, distributed environment.
  • Deep expertise designing, deploying, and operating high-availability, multi-region production architectures on Google Cloud Platform (or AWS/Azure).
  • Strong software engineering background with production-level proficiency in Go, Python, or Java to build distributed automation, internal developer platforms (IDPs), and custom reliability tooling.
  • Deep expertise in building observability solutions using tools such as Datadog, Prometheus, Grafana, or Splunk.
  • Proven ability to define and implement SLIs, SLOs, and alerting strategies.
  • Mastery of Infrastructure as Code (e.g., Terraform, CloudFormation) and CI/CD pipeline automation (e.g., GitHub Actions, Jenkins).
  • Advanced operational experience with Kubernetes, container orchestration, and microservices architectures.
  • Strong background in leading incident response coordination, conducting root-cause analysis, and driving systemic architectural improvements through blameless post-mortems.
  • Demonstrated ability to lead the technical direction of complex, cross-team initiatives, navigating ambiguous challenges and driving scalable solutions from concept to production.
  • Proven track record of mentoring junior and mid-level engineers, fostering technical excellence, and raising the engineering bar through architecture and code reviews.
  • Strong communication skills with the ability to partner effectively across Product, UX, Architecture, Security, and Engineering teams to influence technical direction and prioritize reliability.
  • The knowledge, skills and abilities typically acquired through the completion of a bachelor's degree program or equivalent degree in a field of study related to the job.
  • 3 years of work experience.

Nice To Haves

  • No additional education
  • No additional years of experience
  • None

Responsibilities

  • Develops, tests, deploys, and maintains software, with a clear understanding of the value the software is to provide.
  • Takes a broad view when approaching issues; using a global lens.
  • Consistently achieves results, even under tough circumstances.
  • Develops test suites (functional, destructive, etc) to enable success, rapid deployment of code to production.
  • Takes on new opportunities and tough challenges with a sense of urgency, high energy and enthusiasm.
  • Consistently achieves results, even under tough circumstances.
  • Actively seeks ways to grow and be challenged using both formal and informal development channels.
  • Learns through successful and failed experiment when tackling new problems.
  • Creates new and better ways for the organization to be successful.
  • Delivers multi-mode communications that convey a clear understanding of the unique needs of different audiences.
  • Works the Product Team to ensure user stories are developer ready, easy to understand and testable.
  • Collaborates with other team members in agile processes.
  • Relates openly and comfortably with diverse groups of people.
  • Adapts approach and demeanor in real time to match the shifting demands of different situations.
  • Fields questions from product and engineering teams.
  • Helps grow junior engineers by providing guidance on modern software development frameworks, and leading technical discussions.
  • Notes gaps on the team and provides suggestions for changes to make the team more productive.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service