Reliability Engineer

Global Enterprise Services, LLCFort Meade, MD
Remote

About The Position

Responsible for the availability, performance, monitoring, and incident response, among other things, of the cloud platforms and services. Ensure that everything that goes to production complies with a set of general requirements like diagrams, dependencies of other services, monitoring and logging plans, backups and possible high availability setups. Manages uncaught exceptions, hardware degradation, networking problems, high usage of resources, or slow responses that could happen at any time. Uses metrics such as mean time to recover (MTTR) and mean time to failure (MTTF). Considered an emerging authority, who applies extensive technical expertise. Develops technical solutions to complex problems. Exercises considerable latitude in determining objectives and approaches to assignment.

Requirements

  • 8 years of experience
  • IAT-2 certification
  • 1 or more cloud certifications
  • Bachelor's degree
  • Secret clearance

Responsibilities

  • Ensure availability, performance, monitoring, and incident response of cloud platforms and services.
  • Verify production deployments meet requirements including diagrams, dependencies, monitoring, logging, backups, and high availability.
  • Manage exceptions, hardware degradation, networking problems, high resource usage, and slow responses.
  • Utilize metrics such as Mean Time To Recover (MTTR) and Mean Time To Failure (MTTF).
  • Develop technical solutions to complex problems.
  • Determine objectives and approaches to assignments with considerable latitude.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service