Senior Cloud Site Reliability Engineer (SRE)

Peraton•,
•$104,000 - $166,000•Remote

About The Position

Peraton is looking for a Senior Cloud Site Reliability Engineer (SRE) who will be responsible for designing and developing advanced Python-based AWS cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. This person will apply software engineering practices including Infrastructure-as-Code (IaC) with Terraform to build scalable, reusable solutions and utilities that enhance platform reliability across the Federal Reserve System.

Requirements

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance
  • Bachelors Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience
  • Must have 5+ years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS
  • 7+ years of extensive experience in software development with focus on reliability and platform engineering
  • 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
  • Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices

Nice To Haves

  • Experience with GoLang or additional programming languages is a plus
  • Bachelors Degree in Computer Science, Information Systems, or similar

Responsibilities

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
  • Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
  • Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
  • Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.
  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
  • Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
  • Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios

Benefits

  • Overtime
  • Shift differential
  • Discretionary bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service