Senior Site Reliability Engineer

ECS Tech IncFairfax, VA
$118,000 - $177,000Remote

About The Position

Everforth ECS is seeking a talented Senior Site Reliability Engineer to play a key role in defining, implementing, and growing our SRE practice to ensure the reliability, availability, and performance of our critical production environments. The Senior SRE will contribute to a culture of continuous improvement, identifying areas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency. The successful candidate will have demonstrated hands-on experience designing, implementing, and maintaining solutions to ensure that systems, including infrastructure and applications, are resilient, highly available, and performant. The Senior SRE will also play a critical role in defining and measuring the Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for our solution. The Senior SRE will be responsible for setting up comprehensive logging, monitoring, and alerting solutions using the Elastic stack and other tools as necessary to ensure the continuous performance of services. Additionally, they will respond to incidents, perform root cause analyses, and implement solutions to prevent reoccurrences. The Senior SRE will work in close collaboration with other SRE team members, developers, testers, infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.

Requirements

  • Must be a US citizen with the ability to obtain Public Trust Suitability.
  • 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent.
  • 6+ years of demonstrated experience designing, implementing, and maintaining observability solutions to include logging, monitoring, and alerting.
  • 6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus, Grafana, Splunk, etc.).
  • 3+ years defining and measuring SLOs and SLIs.
  • 3+ years of relevant experience using cloud platforms (AWS GovCloud preferred).
  • 3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.).
  • Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes).
  • Proven ability to collaborate with cross-functional teams (development, testing, and product) to integrate reliability and observability into the software development lifecycle.
  • Strong problem-solving and analytical skills.
  • Proactive, detail-oriented approach to identifying inefficiencies and implementing improvements.
  • Proficient in developing Synthetic monitoring scripts using typescript.

Responsibilities

  • Defining, implementing, and growing the SRE practice.
  • Ensuring the reliability, availability, and performance of critical production environments.
  • Contributing to a culture of continuous improvement.
  • Identifying areas for enhancement and driving initiatives to improve system reliability, scalability, and efficiency.
  • Designing, implementing, and maintaining solutions to ensure systems are resilient, highly available, and performant.
  • Defining and measuring Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Setting up comprehensive logging, monitoring, and alerting solutions using the Elastic stack and other tools.
  • Responding to incidents, performing root cause analyses, and implementing solutions to prevent reoccurrences.
  • Collaborating with SRE team members, developers, testers, infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.

Benefits

  • General Description of Benefits [https://ecstech.com/careers/benefits]
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service