Site Reliability Engineer

Avalore, LLCArlington, VA
$107,000 - $220,000Onsite

About The Position

The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs), identify and resolve performance bottlenecks, and perform root cause analysis on incidents to implement preventative measures and enhance system efficiency. This person will design and implement monitoring and alerting systems to provide visibility into system health and performance, develop and maintain runbooks and playbooks for operational procedures, and automate routine operational tasks to improve efficiency and reduce human error. This person will conduct capacity planning to ensure the system can handle expected loads, implement load testing to validate system performance under stress, and establish disaster recovery procedures to minimize downtime. This person will participate in on-call rotations to respond to system incidents, collaborate with development teams to improve application reliability and performance, and implement chaos engineering practices to identify weaknesses before they impact users.

Requirements

  • SECRET Security Clearance Required
  • Cloud+, GICSP, GSEC, Security+, or SSCP cerification (for Junior / Journeyman roles)
  • SecurityX / CASP+, CCNP Security, CCSP, FITSP-O, or GFACT certification (for Senior role)
  • Good communication skills
  • Self-starter mindset
  • Professional proficiency in English is required
  • Demonstrated proficiency in using all Microsoft Office applications
  • Currently authorized to work in the United States on a full-time basis
  • Bachelor's Degree

Responsibilities

  • Define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs)
  • Identify and resolve performance bottlenecks
  • Perform root cause analysis on incidents to implement preventative measures and enhance system efficiency
  • Design and implement monitoring and alerting systems to provide visibility into system health and performance
  • Develop and maintain runbooks and playbooks for operational procedures
  • Automate routine operational tasks to improve efficiency and reduce human error
  • Conduct capacity planning to ensure the system can handle expected loads
  • Implement load testing to validate system performance under stress
  • Establish disaster recovery procedures to minimize downtime
  • Participate in on-call rotations to respond to system incidents
  • Collaborate with development teams to improve application reliability and performance
  • Implement chaos engineering practices to identify weaknesses before they impact users

Benefits

  • Employer Contributions to Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k, IRA) with a generous matching program
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Short Term & Long Term Disability
  • Training & Development
  • Employee Assistance Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service