Lead Site Reliability Engineer

Peraton•,
•$112,000 - $179,000

About The Position

Peraton is seeking a Lead Site Reliability Engineer to join their team. The role focuses on ensuring the reliability, resilience, and recoverability of mission-essential platforms. This involves leading infrastructure-level disaster recovery drills, platform rebuild validation, and managing automated deployment processes. The engineer will collaborate with cross-functional teams to maintain and enhance complex cloud-based environments, integrate automation solutions, and support modernization and continuity efforts for high-visibility programs.

Requirements

  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma.
  • Expert level hands-on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools.
  • Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.
  • Proven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.
  • Experience building monitoring tools like dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools.
  • Proficiency with programming/scripting languages such as Python, Java, C#, or Go.
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.
  • Strong analytical and documentation skills with the ability to clearly communicate technical findings to cross-functional teams.
  • Ability to work in a fast-paced environment supporting high visibility, mission-critical systems.
  • Ability to obtain a Public Trust clearance.
  • US Citizen or Green Card Holder.

Nice To Haves

  • AWS DevSecOps Engineer certification (preferred).
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and/or Kubernetes certifications (CKA, CKAD).
  • Familiarity with Zero Trust security models and cloud security best practices.
  • Experience with GitLab, Jenkins, or similar CI/CD platforms.
  • Experience with highly regulated environments (healthcare, finance, DHS, DoD, CMS, etc.).
  • Experience supporting federal, defense, or large-scale enterprise programs involving legacy-to-cloud modernization.
  • Prior involvement in large-scale DR drills, continuity of operations (COOP), or portability/executable readiness assessments.

Responsibilities

  • Supporting full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks.
  • Executing infrastructure level drill activities to ensure the platform can be fully rebuilt within the 48-hour recovery target.
  • Verifying end-to-end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations.
  • Identifying exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and driving corrective actions with engineering teams.
  • Designing, implementing, and supporting automated IaaC workflows utilizing Terraform, AWS CloudFormation, and standardized CI/CD pipelines.
  • Managing and optimizing Kubernetes clusters and containerized workloads (Docker), including cluster provisioning, scaling, and workload reliability improvements.
  • Building and maintaining observability solutions using CloudWatch, Datadog, and other monitoring/alerting tools to ensure service reliability and proactive incident response.
  • Developing automation, tooling, and scripts using Python or Java to reduce manual processes and enhance operational repeatability.
  • Collaborating with platform engineering, security, applications, and data teams to ensure consistent, secure, and compliant platform operations.
  • Participating in on-call rotations, root cause analyses, and incident response activities to improve system resilience and operational excellence.

Benefits

  • Overtime
  • Shift differential
  • Discretionary bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service