Senior Site Reliability Engineer (SRE)

ConcentrixUSA Atlanta, GA
Hybrid

About The Position

We are seeking a Senior Site Reliability Engineer (SRE) to build, automate, and support highly scalable cloud-native platforms and digital applications. This role is responsible for improving system reliability, observability, performance, and operational excellence through Infrastructure as Code, Kubernetes, CI/CD automation, monitoring, and incident management. The ideal candidate will have strong experience with AWS/Azure, distributed systems, platform engineering, and enterprise observability tools, with a passion for reducing operational toil and enhancing platform resilience through automation and intelligent operations.

Requirements

  • 5+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Cloud Operations, or a related field.
  • Strong experience with Linux administration, Kubernetes, Docker, and cloud platforms such as AWS and/or Azure.
  • Hands-on experience with Infrastructure as Code (Terraform, CloudFormation, Helm) and CI/CD pipeline automation.
  • Proficiency in scripting and automation using Python and Bash.
  • Experience supporting distributed systems, APIs, microservices, and customer-facing applications in production environments.
  • Strong knowledge of monitoring and observability tools such as Splunk, Grafana, Prometheus, Open Telemetry, and New Relic.
  • Experience with incident management, root cause analysis (RCA), performance tuning, disaster recovery, and reliability engineering.
  • Understanding of Git, DevOps best practices, automation, and cloud-native technologies.
  • Strong problem-solving, troubleshooting, and collaboration skills.
  • Bachelor’s degree in computer science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • Must reside in the United States or have a valid U.S. address for residence.

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient cloud-native platforms and services.
  • Define, monitor, and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
  • Identify, analyze, and eliminate reliability and performance bottlenecks across distributed systems.
  • Lead incident response, Root Cause Analysis (RCA), and implementation of preventive measures.
  • Participate in on-call support and production operations for mission-critical applications and services.
  • Build and manage enterprise observability solutions leveraging tools such as Splunk, Grafana, Prometheus, Datadog, New Relic, and Open Telemetry.
  • Develop dashboards, alerts, monitoring strategies, and reporting frameworks that provide actionable operational insights.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through proactive monitoring and automation.
  • Establish best practices for logging, distributed tracing, application monitoring, and system health management.
  • Develop and enhance shared platform services and infrastructure utilized by multiple engineering teams.
  • Create self-service capabilities and automation tools that improve developer productivity and reduce operational overhead.
  • Design reusable platform components, frameworks, and operational tooling.
  • Collaborate with architecture, engineering, and product teams to define and execute platform modernization strategies.
  • Design, deploy, and manage cloud infrastructure across AWS and Azure environments.
  • Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, Helm, and Kubernetes manifests.
  • Automate infrastructure provisioning, configuration management, deployments, scaling, and recovery processes.
  • Improve infrastructure consistency, governance, security, and scalability through automation and standardization.
  • Build and optimize CI/CD pipelines to enable secure, scalable, and reliable software delivery.
  • Implement deployment automation, release management processes, and validation controls.
  • Support GitOps methodologies and continuous delivery practices.
  • Partner with development teams to improve deployment frequency, quality, and operational stability.
  • Lead operational readiness assessments, disaster recovery planning, and business continuity exercises.
  • Develop and maintain runbooks, playbooks, escalation procedures, and automated remediation workflows.
  • Drive resiliency testing, chaos engineering initiatives, and fault-tolerance improvements.
  • Ensure compliance with operational, security, and reliability standards.
  • Utilize AI-powered observability and incident management platforms to enhance operational efficiency.
  • Leverage predictive analytics and automation to proactively identify risks and prevent service disruptions.
  • Drive adoption of intelligent operational capabilities that improve system reliability, engineering productivity, and customer experience.
  • Explore innovative approaches to reduce manual effort and operational toil through automation and AI-driven solutions.
  • Partner closely with Software Engineering, Platform Engineering, Cloud Architecture, Security, and Product teams to deliver reliable and scalable solutions.
  • Provide technical leadership and guidance on reliability best practices, automation strategies, and operational excellence.
  • Mentor junior engineers and contribute to continuous improvement initiatives across the organization.

Benefits

  • medical, dental, and vision insurance
  • comprehensive employee assistance program
  • 401(k) retirement plan
  • paid time off and holidays
  • paid learning days
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service