Director, Sustaining Engineering - Spark

Crusoe•Denver, CO
•$225,000 - $255,000

About The Position

Crusoe is a vertically integrated AI Factory company with a mission to accelerate the abundance of energy and intelligence. Our competitive advantage, "Speed is the only moat," is directly tied to our ability to rapidly design, manufacture, and deploy our own modular power and compute infrastructure. We are seeking a Sustaining Engineering Leader to own how Spark performs after it is deployed. Product Management defines the next generation. New Product Introduction delivers it into production. You own each generation from the moment it leaves ramp: fleet reliability, field failures, configuration, obsolescence, and the design changes that keep deployed units performing for their full service life. You do not set the roadmap and you do not run the factory. You own the installed base, and you are the evidence source Product Management and NPI both depend on. You will start as the single-threaded owner of this function and build the team as the fleet grows. Spark deployments begin in early 2027. You will build the system before there is a fleet to run it on. Stand up FRACAS, the root cause evidence standard, and the Failure Review Board. Establish the as-built configuration baseline before there are units to reconcile. Define spares, service intervals, and repair versus replace policy for the first generation. Participate in the production release and ramp readiness reviews for the first deploying generation as the receiving organization.

Requirements

  • Bachelor's degree in Mechanical, Electrical, Reliability, or a closely related engineering discipline.
  • 10 to 15 years in sustaining, reliability, or product support engineering for deployed capital equipment or infrastructure hardware.
  • Direct ownership of a fielded installed base, with field data converted into design change that measurably reduced failure rate or service cost.
  • Obsolescence and lifecycle management on a product whose service life exceeds that of the components inside it.
  • Experience standing a function up from nothing, including the processes and the partner agreements behind it.
  • FRACAS, structured root cause analysis, and mean time between failures and population failure rate analysis.
  • PLM, change order workflow, as-built control, and serial level effectivity.
  • DMSMS practice, end of life monitoring, last time buy modeling, and alternate part qualification.
  • Electrical, mechanical, and thermal command sufficient to adjudicate root cause on a modular power and compute product, including liquid cooling.
  • Spares, service cost per unit, and the link between availability and revenue.

Nice To Haves

  • Experience with modular, containerized, or prefabricated infrastructure products in the field.
  • Familiarity with liquid cooled data center hardware, including CDU and rack level cooling interfaces.
  • The persistence to chase a root cause past the easy answer, and the judgment to know when a fleet-wide fix is worth its disruption.

Responsibilities

  • Take technical ownership of each generation once it completes production ramp, through a defined handover from New Product Introduction covering the released configuration, open issues, and known risks.
  • Own failure rate by subsystem, mean time between failures, and repeat failure rate, and define the telemetry every unit must report to make them measurable.
  • Run FRACAS and chair the Failure Review Board. Close every failure with a root cause, a verified fix, and proof it worked.
  • Contain first for safety or fleet-wide risk, and settle attribution after. Own the retrofit decision and its sequencing.
  • Own engineering change for generations in the field, with cost and fleet impact quantified before approval.
  • Own the configuration baseline and define what Operations records. Reconcile as-maintained against baseline so a change targets exactly the units that need it.
  • Run a proactive DMSMS program. Model last time buys against fleet demand and qualify alternates before a shortage forces the choice.
  • Issue engineering change packages, kit definitions, preventive maintenance intervals, and acceptance criteria.
  • Convert fleet evidence into quantified design requirements. Product Management decides what enters the PRD.
  • Supply field-derived requirements and acceptance criteria, so a failure the fleet has already seen must be designed out and proven before the next generation is released.
  • Own the reliability and service inputs to Spark unit economics.
  • Represent serviceability and maintainability in design reviews. You hold no approval authority; you make the case with fleet data.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service