Site Reliability Manager

Karsun Solutions, LLCHerndon, VA
Onsite

About The Position

We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability, scalability, and performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate must reside in DMV area and be available to work on site in office or customer locations in this area. Must have demonstrated experience in having performed this role for at least 3 years.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field; Master's degree preferred.
  • 10+ years of experience in a similar role managing a team of site reliability engineers and delivering in the AWS cloud platform.
  • 5+ years of experience supporting operations and maintenance for cloud-native applications in production that are fault-tolerant, self-healing, scalable, and highly available.
  • Deep understanding of the AWS cloud computing platform and containerization technologies (e.g., Docker, Kubernetes).
  • Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible, ArgoCD) and CI/CD pipelines.
  • Experience with Datadog as the primary logging, monitoring, and observability platform (alongside AWS Cloudwatch).
  • Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams.
  • Strong problem-solving and analytical skills, with a keen attention to detail.
  • Ability to obtain and maintain a Public Trust clearance.
  • Must reside in DMV area and be available to work on site in office or customer locations in this area.
  • Must have demonstrated experience in having performed this role for at least 3 years.

Nice To Haves

  • Certifications such as AWS Certified DevOps Engineer are a plus.
  • Understanding of modern architecture, e.g., micro-services, EDA, etc., and a cautious approach against overcomplexity and overengineering.
  • Experience designing and operating distributed systems and cloud infrastructure at scale.

Responsibilities

  • Lead an 8-20 person service delivery team (Service Support specialist, DevSecOps, and Site Reliability engineers), mentoring them to foster a culture of learning and innovation.
  • Take joint ownership of production reliability, standardize observability and error handling using Datadog, develop SLOs and KPIs to measure system performance, and conduct incident post-mortems and root cause analyses.
  • Define and implement best practices for infrastructure as code, deployment automation, and drive vulnerability management and resolution.
  • Oversee the end-to-end platform lifecycle, collaborate with cross-functional teams to design scalable and fault-tolerant architectures, and drive continuous improvement initiatives to enhance system efficiency.

Benefits

  • All qualified applicants will receive consideration for employment without regard to disability, status as a protected veteran or any other status protected by applicable federal, state, local, or international law.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service