About The Position

Nscale is seeking a Design Reliability Engineer to join their Power and Energy group. This role is responsible for ensuring systems meet availability targets through rigorous analysis and disciplined engineering practices. The Power and Energy group develops behind-the-meter generation, including reciprocating engines, battery energy storage, and medium/high-voltage distribution, operating as islanded microgrids. Availability is a critical, engineered quantity that must be allocated, modeled, defended against risks, and measured live. This role will own this process end-to-end, integrating power and mechanical scopes. It is a hands-on technical leadership position requiring minimal oversight. The engineer will own the availability model of record, the FMEA program, and the reliability foundations of the digital twin. They will direct consultants, challenge methodologies, and present analysis to executives. The role sits within the Power and Energy technology organization, collaborating with electrical, controls, mechanical, and pipeline engineering managers, as well as commissioning and operations teams. The engineer will act as the owner's technical authority on reliability, directing consultant work and resolving complex reliability questions for large, redundant systems.

Requirements

  • 5+ years of reliability engineering experience on power generation, process, or mission-critical facilities.
  • Significant experience owning RAM analysis for large, redundant systems.
  • Deep RAM modeling capability across discrete-event simulation (BlockSim, Raptor, or equivalent) and closed-form analytical methods.
  • Proven FMEA/FMECA leadership on major rotating, electrical, or process equipment.
  • Fluency with standard failure data sources (IEEE 493, OREDA, IEEE 3006 series or similar) and their limitations.
  • Experience allocating availability targets from commercial commitments down to systems.
  • Experience defending reliability analysis to executives, customers, or insurers.
  • Sharp eye for common-mode and dependent failure mechanisms.
  • Track record of identifying risks that redundancy arithmetic hides.
  • Working software capability: comfort scripting analyses (Python or similar) and structuring reliability data for automation.
  • Experience feeding reliability analysis into live design processes on major capital projects.
  • Demonstrated ability to run a scope with minimal oversight.
  • Bachelor's degree in Mechanical, Electrical, Chemical, or Reliability Engineering or a related field.
  • Willing and able to travel to project sites and vendor facilities regularly (typically 15–25%, higher during commissioning campaigns).

Nice To Haves

  • Exposure to digital twin, telemetry-driven reliability, or operational availability measurement programs.
  • CRE certification or equivalent.
  • PE license.

Responsibilities

  • Own the availability model of record for each generation campus, spanning generation, electrical distribution, fuel supply, and cooling.
  • Allocate committed SLA targets down through the system, defining availability budgets, redundancy requirements, MTTR assumptions, and sparing/maintenance strategies.
  • Apply appropriate RAM modeling methods (e.g., Monte Carlo simulation, closed-form analytical models) and defend the chosen methodology.
  • Identify and mitigate common-mode and dependent failures.
  • Report availability in both single-path and contracted-capacity views, including sensitivity analyses.
  • Lead the FMEA/FMECA program for major equipment and systems (engines, generators, BESS, switchgear, transformers, fuel gas systems, cooling infrastructure).
  • Curate failure rate and repair data from industry sources, OEMs, and field operations.
  • Influence design decisions through reliability analysis, including redundancy configuration, single-point-of-failure treatment, equipment selection, and maintainability requirements.
  • Integrate reliability analysis into design reviews, HAZOPs, and vendor evaluations.
  • Own the reliability core of Nscale's digital twin, including model architecture, data structures, and analytical methods for real-time availability computation.
  • Define operational data requirements for the digital twin, ensuring trustworthy data input.
  • Collaborate with software and data teams to transition RAM analysis from static studies to automated, queryable tooling.
  • Establish feedback loops for measured availability versus model prediction, model recalibration, and reliability growth tracking.
  • Support commissioning and startup with reliability input, including burn-in and reliability run design, failure tracking, and acceptance criteria.
  • Lead and support root cause analyses of significant failures and availability events, driving corrective actions.
  • Develop Nscale's reliability engineering standards, methods, and reference models.
  • Direct RAM and reliability consultants, setting work basis, reviewing deliverables, and integrating their output.
  • Coordinate with data center reliability and operations teams to ensure availability is engineered and measured across the full path to the GPU.
  • Grow internal reliability engineering capability by building a team as the portfolio scales.
  • Communicate complex reliability analysis clearly to project leadership and executives, including assessments of confidence, sensitivity, and risk.

Benefits

  • Highly competitive US compensation package (base + bonus + equity)
  • Performance reviews every 12 months
  • Comprehensive medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Generous PTO plus US federal holidays
  • A career-defining opportunity to be an early member of one of the fastest growing AI infrastructure companies in the world
  • Flexible workplace
  • Autonomy to get the job done
  • Medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service