Remote Operational Resilience Engineer (VA ESOM)

KentroUNAVAILABLE, UNAVAILABLE
Remote

About The Position

Kentro is hiring for a Remote Operational Resilience Engineer to support VA ESOM contract. The Engineer will identify non-cyber technology and operational risks that could prevent key applications from meeting recovery objectives. This role evaluates application architecture, infrastructure dependencies, recovery procedures, staffing, vendors, and other resources through tabletop exercises and related technical reviews. The engineer works with system owners to define and validate corrective actions. This position can be performed remotely within the United States and will support Eastern Time working hours.

Requirements

  • Master’s degree in computer science, information systems, engineering, or another relevant technical discipline.
  • 10 years of relevant experience.
  • 10 years of additional relevant experience may be substituted for education.
  • Experience in systems engineering, application architecture, infrastructure resilience, site reliability, disaster recovery, or a related discipline.
  • Knowledge of cloud and on-premises platforms, networking, databases, storage, backup, replication, failover, and application integration.
  • Experience analyzing application dependencies, recovery sequences, and technical procedures.
  • Understanding of recovery time objectives, recovery point objectives, restoration priorities, and business impact information.
  • Strong troubleshooting, facilitation, documentation, and communication skills.
  • Excellent communication and collaboration skills, with the ability to work across engineering, operations, and program teams.
  • US Citizen or Lawful Permanent Resident (Green Card)
  • Willing and able to obtain and maintain Public Trust Clearance
  • Must meet updated ID requirements

Nice To Haves

  • Experience applying NIST SP 800-34, NIST SP 800-84, or NIST SP 800-53 contingency planning controls.
  • Experience leading application failover, backup restoration, or full recovery tests.
  • Knowledge of high-availability design, infrastructure as code, automation, and recovery orchestration.
  • Relevant cloud, systems, disaster recovery, or business continuity certification.
  • Background supporting large federal, DoD, or public-sector IT environments.

Responsibilities

  • Identify technology, process, staffing, facility, vendor, and dependency risks that could affect application recovery.
  • Review application architectures, data flows, interfaces, infrastructure dependencies, and restoration sequences.
  • Assess whether recovery strategies support approved recovery time and recovery point objectives.
  • Evaluate backup, restoration, alternate processing, failover, reconstitution, communications, and return-to-normal procedures.
  • Develop tabletop scenarios involving infrastructure failure, service loss, unavailable facilities, vendor outages, capacity constraints, and data restoration issues.
  • Verify that plans identify recovery roles, decision authorities, escalation paths, required resources, and external dependencies.
  • Facilitate technical discussions and evaluate participant actions against exercise objectives.
  • Document weaknesses, supporting evidence, root causes, and unresolved assumptions.
  • Develop corrective actions and recommend hands-on tests when discussion alone cannot validate recovery capability.
  • Support follow-up reviews that confirm procedures and engineering changes address the original finding.

Benefits

  • paid time off
  • healthcare benefits
  • supplemental benefits
  • 401k including an employer match
  • discount perks
  • rewards
  • education reimbursement for certifications, degrees, or professional development
  • funds for activities – virtual and in-person
  • happy hours
  • holiday events
  • fitness & wellness events
  • annual celebrations
  • charity galas/events
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service