Design Reliability Engineer — Power and Energy

Nscale•Houston, TX
•$130,000 - $200,000•Hybrid

About The Position

Nscale is building the infrastructure for the AI revolution, developing cutting-edge, sovereign generative AI solutions powered by high-performance, sustainable data centers and GPUs. The rapid growth of artificial intelligence is driving unprecedented global demand for compute and GPU infrastructure. Nscale is delivering the platforms that will enable the next decade of innovation by integrating next-generation compute hardware and GPU clusters. This role is a hands-on technical leadership position within the Power and Energy group, which develops behind-the-meter generation colocated with data center campuses. This includes fleets of reciprocating engines, battery energy storage, and medium- and high-voltage distribution operating as islanded microgrids. The role owns the availability model of record for each campus, the FMEA program across major equipment, and the reliability foundations of the digital twin. The engineer will direct consultants and reliability data work, challenge methodology, and present analysis to executives. This position sits within the Power and Energy technology organization, collaborating with electrical, controls, mechanical, and pipeline engineering managers, as well as commissioning and operations teams. The role requires an owner's technical authority on reliability, resolving complex system challenges.

Requirements

  • 5+ years of reliability engineering experience on power generation, process, or mission-critical facilities, with significant time owning RAM analysis for large, redundant systems.
  • Deep RAM modeling capability across both discrete-event simulation (BlockSim, Raptor, or equivalent) and closed-form analytical methods, with the judgment to know which the question deserves.
  • Proven FMEA/FMECA leadership on major rotating, electrical, or process equipment, and fluency with the standard failure data sources (IEEE 493, OREDA, IEEE 3006 series or similar) and their limitations.
  • Experience allocating availability targets from commercial commitments down to systems, and defending the resulting analysis to executives, customers, or insurers.
  • A sharp eye for common-mode and dependent failure mechanisms, and a track record of finding the risks that redundancy arithmetic hides.
  • Working software capability: comfort scripting analyses (Python or similar) and structuring reliability data for automation, not just operating desktop tools.
  • Experience feeding reliability analysis into live design processes on major capital projects, and the standing to influence engineering decisions with it.
  • Demonstrated ability to run a scope with minimal oversight: establishing the basis, setting the pace, and escalating with solutions rather than problems.
  • Bachelor's degree in Mechanical, Electrical, Chemical, or Reliability Engineering or a related field.
  • Willing and able to travel to project sites and vendor facilities regularly (typically 15–25%, higher during commissioning campaigns).

Nice To Haves

  • Exposure to digital twin, telemetry-driven reliability, or operational availability measurement programs is a strong advantage.
  • CRE certification or equivalent.
  • PE license preferred.

Responsibilities

  • Own the availability model of record for each generation campus: system-level RAM models spanning generation, electrical distribution, fuel supply, and cooling, built and maintained as living engineering assets.
  • Allocate committed SLA targets down through the system: availability budgets by subsystem, redundancy requirements, MTTR assumptions, and sparing and maintenance strategies that make the top-level number achievable.
  • Apply the right method for each question — Monte Carlo simulation (BlockSim or equivalent) where it earns its complexity, closed-form analytical models where they are faster and more auditable — and defend the choice on its merits.
  • Hunt common-mode and dependent failures relentlessly: shared fuel supply, shared cooling, control system dependencies, and site-wide events that redundancy counts conceal.
  • Report availability the way commitments are written: both single-path and contracted-capacity views, with sensitivities that show which assumptions carry the number.
  • Own the FMEA/FMECA program for major equipment and systems: engines and generators, BESS, switchgear and transformers, fuel gas systems, and cooling infrastructure, establishing baselines and keeping them current as designs mature.
  • Curate the failure rate and repair data underpinning every model: industry sources such as IEEE 493 and OREDA, OEM data challenged rather than transcribed, and field data as it accumulates.
  • Turn analysis into design influence: redundancy configuration, single-point-of-failure treatment, equipment selection input, and testability and maintainability requirements fed into the engineering teams while designs can still change.
  • Bring reliability analysis into design reviews, HAZOPs, and vendor evaluations as a routine discipline, not a report delivered after decisions are made.
  • Own the reliability core of Nscale's digital twin: the model architecture, data structures, and analytical methods that let availability be recomputed on the fly as system state, configuration, and failure data change.
  • Define the operational data requirements — event capture, failure coding, downtime attribution, and run-hour tracking — so the twin is fed by trustworthy data from day one of operations.
  • Work with our software and data teams to move RAM analysis from static studies into instrumented, queryable tooling, and champion analytical approaches that are automatable and auditable rather than locked in desktop tools.
  • Establish the feedback loop: measured availability versus model prediction, model recalibration, and reliability growth tracking from commissioning onward.
  • Support commissioning and startup with reliability input: burn-in and reliability run design, failure tracking during startup, and acceptance criteria grounded in the availability model.
  • Lead and support root cause analyses of significant failures and availability events, and drive corrective actions back into designs, models, and standards.
  • Develop Nscale's reliability engineering standards, methods, and reference models, so each campus builds on the last instead of starting over.
  • Direct RAM and reliability consultants: set the basis they work to, review and challenge their methodology and deliverables, and integrate their output into one coherent picture across power and mechanical scopes.
  • Coordinate with data center reliability and operations teams so that availability is engineered and measured across the full path to the GPU, not just to the fence line.
  • Grow the internal reliability engineering capability over time, building a small team as the portfolio scales.
  • Communicate complex reliability analysis clearly to project leadership and executives, with honest assessments of confidence, sensitivity, and risk.

Benefits

  • Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months
  • Comprehensive medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Generous PTO plus US federal holidays
  • A career-defining opportunity to be an early member of one of the fastest growing AI infrastructure companies in the world
  • A human-first approach: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy needed to get the job done.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service