Compute Deployment Engineer

FluidstackSan Francisco, CA
$150,000 - $250,000Hybrid

About The Position

Fluidstack is building civilization-scale infrastructure for AI, aiming to deploy gigawatts of compute infrastructure faster than anyone else. This role is on the Infrastructure Team, which is responsible for bringing gigawatts of accelerators from first power-on to production, ensuring facility availability and rapid rack qualification. The team focuses on scaling through tooling rather than headcount, with a goal of growing deployed megawatts significantly while keeping the team size stable.

Requirements

  • Experience bringing up server or GPU fleets at scale (hundreds of nodes or more) and taking them to production.
  • Deep proficiency in Linux and out-of-band management tools such as BMC, IPMI, and Redfish.
  • Experience automating hardware workflows using Python or Go.
  • Experience working physically in data halls, including racking, cabling, and component swapping.
  • Ability to act as remote hands or direct remote hands effectively.
  • Methodical approach to triaging failures across hardware, firmware, and software to isolate faults to a component.
  • Willingness to travel for turn-up windows when new data halls come online.

Nice To Haves

  • Kubernetes-based bare-metal provisioning experience.
  • Accelerator platform bring-up experience (NVIDIA, AMD, or custom).
  • Experience with burn-in and stress harness design.
  • Experience with DCIM and inventory tooling.

Responsibilities

  • Own compute turn-up from facility availability to ready-for-service, the period after network handoff and before customer workloads.
  • Qualify racks at scale, including establishing firmware baselines, configuring BMC and BIOS, running burn-in tests, and performing node and cluster level validation on GPU and custom accelerator platforms.
  • Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), automating runs to reduce qualification queues.
  • Triage hardware failures identified during qualification, isolating them to specific components, managing RMAs and vendor escalations, and feeding failure patterns back into qualification processes.
  • Perform turn-up remotely by default, with on-site visits of approximately one week per data hall as new halls become facility-available, plus occasional overlapping-site weeks.
  • Collaborate with network deployment, ICT, data center operations, and hardware teams during turn-up windows.
  • Support incident response for newly live capacity.

Benefits

  • Commitment to pay equity and transparency.
  • Equal Employment Opportunity Employer status.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service