SRE - DevOps Engineer

Capgemini•Dallas, TX
•$86,129 - $127,189•Onsite

About The Position

We are seeking a highly motivated Site Reliability Engineer to help build and operate reliable, scalable, and secure services across our platform. This role is designed for someone who combines strong DevOps practices, modern SRE principles, and software engineering experience to improve system reliability, automate operations, and support high-availability production environments. The ideal candidate will be passionate about building resilient systems, improving developer productivity, and driving operational excellence through automation, observability, and engineering best practices. This role partners closely with engineering, platform, and product teams to ensure services are built for reliability from the start and remain performant, stable, and supportable at scale.

Requirements

  • Key Skills - Node.js, Python, DevOps, Jenkins, AWS

Nice To Haves

  • 9+ years of demonstrated experience developing and designing products around Site Reliability Engineering principles to improve stability and platform availability for containerized workloads and on-premises services using Kubernetes.
  • Experience managing and interpreting large datasets using query languages and creating dashboards and reports with Power BI and Grafana.
  • Strong background in managing cloud and on-premises infrastructure using Infrastructure as Code tools, including Terraform, and CloudFormation.
  • Hands-on experience building, operating, monitoring, logging, and alerting distributed systems at scale using Datadog and Splunk.
  • Experience supporting DevOps practices for service delivery and operations using Jenkins, Azure DevOps, Team Foundation Version Control, and CI/CD automation.
  • Experience developing software and automation solutions to support application delivery, operations, and repeatable business processes using Python.
  • Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway
  • Strong development experience in scripting, automation, and integration across Linux and Windows-based environments.

Responsibilities

  • Design, build and operate resilient, scalable systems using DevOps, SRE, and software development best practices.
  • Deliver high-availability services through automation, infrastructure as code, and proactive reliability engineering.
  • Improve monitoring, logging, alerting, and observability for distributed systems.
  • Support CI/CD automation, deployment workflows, and production tooling to reduce operational toil.
  • Drive incident response, root cause analysis, and recovery improvements to minimize downtime.
  • Partner with engineering teams to embed reliability into the software development lifecycle.
  • Automate provisioning, configuration, and self-healing across cloud and on-prem environments.
  • Validate resiliency and performance through testing, chaos engineering, and capacity planning.

Benefits

  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave
  • Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
  • Other benefits as provided by local policy and eligibility
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service