Senior SRE Engineer/DevOps Engineer

Lumen Solutions Group IncWashington, DC
Remote

About The Position

We are seeking a hands-on Senior Site Reliability Engineer (SRE) / DevOps Engineer to support cloud infrastructure, automation, and production operations. The ideal candidate will have strong expertise in AWS, Infrastructure as Code (IaC), CI/CD automation, observability, and incident management. This role focuses on improving system reliability, automating operational processes, supporting production environments, and implementing SRE best practices to ensure highly available and scalable applications.

Requirements

  • 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Infrastructure Engineering.
  • Strong experience with AWS cloud services (Azure experience is a plus).
  • Hands-on experience with GitHub Actions, Jenkins, and AWS CodePipeline.
  • Expertise in Terraform, CloudFormation, or AWS CDK.
  • Strong scripting skills in Python.
  • Experience with Ansible or other configuration management tools.
  • Experience with Dynatrace observability and application monitoring.
  • Strong knowledge of Docker, Kubernetes, and Amazon ECS.
  • Solid understanding of Linux administration, networking, and cloud infrastructure.
  • Experience with ServiceNow, ITIL processes, incident management, and production support.
  • Knowledge of relational and NoSQL databases.
  • Excellent troubleshooting, communication, and documentation skills.
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).

Nice To Haves

  • Experience supporting enterprise-scale cloud environments.
  • Knowledge of performance testing, resiliency engineering, and operational excellence.
  • Experience implementing automation and self-service operational tools.
  • Familiarity with cloud cost optimization and reliability engineering practices.

Responsibilities

  • Design, build, and maintain CI/CD pipelines using GitHub Actions, Jenkins, and AWS CodePipeline.
  • Automate infrastructure provisioning using Terraform, CloudFormation, or AWS CDK.
  • Develop automation tools and scripts using Python and configuration management tools such as Ansible.
  • Deploy, configure, and optimize Dynatrace for monitoring, distributed tracing, dashboards, alerting, and anomaly detection.
  • Support AWS cloud infrastructure and production environments, including participation in on-call rotations.
  • Perform incident response, troubleshooting, root cause analysis (RCA), and create knowledge base documentation.
  • Implement and maintain SRE best practices, including SLIs, SLOs, error budgets, resiliency testing, and performance monitoring.
  • Configure auto-scaling, capacity planning, and operational cost optimization.
  • Manage Linux-based environments, containers (Docker, Kubernetes, ECS), networking, and cloud infrastructure.
  • Support security initiatives, including access management, certificate management, and remediation of security incidents.
  • Collaborate with cross-functional engineering teams to improve platform reliability, scalability, and operational efficiency.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service