Reliability Engineer

Applied Information SciencesReston, VA

About The Position

As your initial project assignment, you will support the unique needs of our client as a Reliability Engineer. AIS is supporting a federal customer with a program focused on architecture and infrastructure. This role is responsible for the availability, performance, monitoring, and incident response of cloud platforms and services. The engineer will ensure that all production deployments comply with general requirements such as diagrams, service dependencies, monitoring and logging plans, backups, and high availability setups. They will manage uncaught exceptions, hardware degradation, networking problems, high resource usage, or slow responses. The role utilizes metrics like mean time to recover (MTTR) and mean time to failure (MTTF). This position requires extensive technical expertise and the ability to develop technical solutions to complex problems, with considerable latitude in determining objectives and approaches.

Requirements

  • Bachelors degree in Computer Science, Information Systems, Engineering, or related field (or equivalent experience).
  • 8+ years of relevant experience supporting enterprise cloud and/or infrastructure environments.
  • Certifications: IAT-2, 1 or more cloud certifications.
  • Active Secret clearance (or higher).
  • Experience working in regulated environments and following secure engineering / documentation practices.

Nice To Haves

  • Experience supporting DoD/IC programs and mission systems.

Responsibilities

  • Responsible for the availability, performance, monitoring, and incident response, among other things, of the cloud platforms and services.
  • Ensure that everything that goes to production complies with a set of general requirements like diagrams, dependencies of other services, monitoring and logging plans, backups and possible high availability setups.
  • Manages uncaught exceptions, hardware degradation, networking problems, high usage of resources, or slow responses that could happen at any time.
  • Uses metrics such as mean time to recover (MTTR) and mean time to failure (MTTF).
  • Develops technical solutions to complex problems.
  • Exercises considerable latitude in determining objectives and approaches to assignment.

Benefits

  • Employee Ownership
  • Continuous Learning
  • Inclusive Culture
  • Mission-Driven Work
  • Competitive and fair compensation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service