SRE-Azure - Global Industrial

Genuine Parts Company•Atlanta, AL

About The Position

The Azure Site Reliability Engineer (SRE) is responsible for improving the reliability, availability, scalability, performance, and operational excellence of enterprise applications and platforms hosted within Microsoft Azure. This role combines software engineering, cloud infrastructure, automation, and DevOps practices to build and support resilient, cloud-native solutions while reducing operational toil through automation. The Azure SRE leverages Azure platform services, Kubernetes, Infrastructure as Code (IaC), and observability tools to ensure mission-critical systems remain highly available, secure, and performant. This role partners closely with application development, cloud engineering, architecture, and cybersecurity teams to drive continuous improvement, accelerate cloud adoption, and maintain service reliability through proactive monitoring, incident management, and operational engineering.

Requirements

  • Typically requires a bachelor's degree and five (5) to seven (7) years of experience in a technology and/or software engineering role or an equivalent combination.
  • Experience supporting large-scale, highly available, distributed applications and cloud platforms.
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, and Operational Excellence.
  • Hands-on experience with Microsoft Azure services, including infrastructure, networking, security, platform services, and cloud-native technologies.
  • Experience administering and supporting Azure Kubernetes Service (AKS), Kubernetes clusters, containers, and scalable distributed systems.
  • Proficiency with Infrastructure as Code (Terraform, ARM Templates) and Git-based deployment practices.
  • Experience with Azure DevOps, GitHub Actions, CI/CD pipelines, release automation, and DevOps methodologies.
  • Strong troubleshooting skills across cloud infrastructure, operating systems, databases, networking, and security domains.
  • Experience with monitoring and observability platforms, including Azure Monitor, Log Analytics, Application Insights, Grafana, Datadog, and Dynatrace.
  • Knowledge of microservices, APIs, distributed architectures, and cloud-native design patterns.
  • Working knowledge of Windows Server, Linux, networking, DNS, load balancing, firewalls, and hybrid cloud connectivity.
  • Experience with capacity planning, performance engineering, scalability testing, disaster recovery, and business continuity practices.
  • Strong analytical, problem-solving, communication, and collaboration skills.

Responsibilities

  • Monitor, analyze, and optimize system performance, availability, and reliability across Azure-hosted platforms and applications.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
  • Drive continuous service improvement through operational metrics, trend analysis, reliability engineering practices, and platform modernization efforts.
  • Partner with development teams to improve service reliability through testing, release validation, deployment automation, and production readiness reviews.
  • Design, build, and support Azure infrastructure and platform services, including Azure Kubernetes Service (AKS), Azure Networking, App Services, and Storage
  • Develop and maintain Infrastructure as Code (IaC) solutions utilizing Terraform and Azure-native deployment technologies.
  • Lead and support incident response, root cause analysis (RCA), post-incident reviews, and service restoration efforts.
  • Automate operational processes, platform provisioning, deployments, and remediation activities to reduce manual effort (TOIL) and improve reliability.
  • Identify, investigate, and mitigate performance, security, networking, and availability issues, including traffic anomalies and service disruptions.
  • Participate in on-call rotations and maintain operational documentation, runbooks, and knowledge articles as needed.

Benefits

  • options for healthcare coverage
  • 401(k)
  • tuition reimbursement
  • vacation
  • sick
  • holiday pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service