Fleet Reliability Engineer

QuartermasterArlington, VA
Hybrid

About The Position

Quartermaster is developing a real-time ocean monitoring network using civil and commercial vessels as data collection points. Their SmartMast™ system, a ruggedized sensor package, collects HD video, signal intelligence, and AI-powered insights. The company is seeking its first dedicated Fleet Reliability Engineer to ensure the health, uptime, and reliability of this deployed hardware fleet. This role involves monitoring fleet telemetry, performing root-cause analysis on failures, and collaborating with design, firmware, and field-service teams to prevent recurring issues. It's a hands-on position for a senior engineer experienced in hardware, embedded systems, data analysis, and field operations, tasked with building a reliability program from the ground up, including metrics, monitoring, failure tracking, and maintenance workflows for a scaling fleet. The Fleet Reliability Engineer will be responsible for fleet-wide reliability metrics such as uptime, availability, MTBF, MTTR, and data-yield. This includes developing dashboards and alerts to identify degrading units before they fail, defining system health parameters, and leading root-cause analysis for field failures. The engineer will maintain a failure database, drive FMEA and CAPA processes, and translate field insights into design improvements. Additionally, the role involves defining maintenance schedules, spares strategy, and RMA processes, as well as updating field procedures and directly supporting complex repairs, potentially involving travel. The engineer will also collaborate with the hardware team on environmental and life-testing protocols, establish reliability requirements for new hardware revisions, and ensure reliability processes are scalable for a growing fleet.

Requirements

  • Bachelor's degree in Electrical, Mechanical, Systems, Reliability, or a related engineering discipline — or equivalent hands-on experience.
  • 5+ years of engineering experience with deployed electro-mechanical hardware, at least 2 of which are in reliability, sustaining/field engineering, or hardware operations for a fielded product.
  • Demonstrated ownership of hardware reliability outcomes for a fleet or installed base — you have been directly responsible for uptime, failure rates, or MTBF/MTTR of real hardware in the field.
  • Hands-on proficiency with root-cause analysis methods (8D, 5-Whys, fishbone) and reliability tools such as FMEA, fault-tree analysis, and CAPA.
  • Practical experience diagnosing electro-mechanical systems using telemetry/logs, bench instruments, and physical teardown.
  • Data fluency: able to query, analyze, and visualize fleet telemetry using SQL and Python (or equivalent) to find trends and drive decisions.
  • Working knowledge of electronics, power systems, and mechanical enclosures, and the failure modes of hardware operating in harsh outdoor environments.
  • Willingness and ability to travel periodically to field sites (vessels, ports, installation locations), including occasional international travel.
  • Must be legally authorized to work in the United States and able to satisfy any customer- or contract-driven eligibility requirements associated with government and maritime-security work.

Nice To Haves

  • Experience with hardware deployed in marine, maritime, offshore, automotive, aerospace/defense, satellite, telecom, or other remote/harsh-environment fleets.
  • Familiarity with IP-rated enclosures, corrosion and salt-fog effects, marine power systems, and environmental qualification (e.g., IEC 60529, MIL-STD-810, IEC 60068).
  • Experience with connected/IoT or edge devices: remote diagnostics, OTA firmware updates, and interpreting embedded-system and connectivity (SATCOM/cellular) telemetry.
  • Exposure to camera/optical systems, RF/software-defined radios, batteries, or edge-AI compute hardware.
  • Background building a reliability or sustaining-engineering function from scratch at a hardware startup or scaling operation.
  • ASQ Certified Reliability Engineer (CRE) or comparable credential.

Responsibilities

  • Own fleet-wide reliability metrics — uptime, availability, MTBF, MTTR, data-yield, and failure rates by component and by deployment environment — and report them to engineering and leadership.
  • Build and refine dashboards, alerting, and telemetry pipelines that surface degrading units (power, thermal, connectivity, camera, radio, compute) before they go offline.
  • Define what "healthy" means for each subsystem and set the thresholds that trigger proactive intervention.
  • Lead root-cause analysis (RCA) on field failures, from telemetry forensics through physical teardown of returned units.
  • Maintain the fleet failure database and drive FMEA, reliability growth tracking, and corrective/preventive action (CAPA) to closure.
  • Close the loop with hardware, firmware, and manufacturing teams — translating field failures into design-for-reliability, component-selection, and firmware changes.
  • Define preventive maintenance schedules, spares strategy, and the RMA / repair-and-return process for a globally distributed fleet.
  • Update installation, diagnostic, and field-repair procedures and troubleshooting guides used by internal technicians and partner crews.
  • Support field deployments and complex repairs directly, including periodic travel to vessels, ports, and installation sites.
  • Work with the hardware team to establish environmental and life-test protocols (vibration, salt-fog/corrosion, thermal, ingress, power) to qualify hardware and predict field life before deployment.
  • Feed reliability requirements and acceptance criteria into new hardware revisions and supplier qualification.
  • Design the reliability processes and tooling so they scale as the fleet grows into the thousands of units.

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • 401k
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service