Product Manager, Compute Operations

FluidstackAustin, TX
$245,000 - $300,000Onsite

About The Position

Fluidstack is building civilization-scale infrastructure for AI, focused on delivering 10 to 100s of GWs of compute faster than anyone else. This role is for a Product Manager to own the automation roadmap for compute production. The company operates with a philosophy of full autonomy, insane urgency, reasoning from first principles, and a deep care for the problem space. The Decision Team is responsible for automating the delivery of gigawatts, forward-deploying beside experts, and delivering supercomputers faster than the last. This role will be the first product manager for compute production automation, deciding which aspects of GPU fleet health, repair, hardware qualification, on-call, and facility maintenance become software. The goal is to automate processes that currently rely on manual tracking and intervention, using software to manage schedules, decisions, and tasks generated from a live knowledge graph. The product manager will also be responsible for landing systems already in flight, such as maintenance systems, asset registers, and retiring legacy datacenter inventory. A key aspect of the role involves embedding with production engineers and facility operators, participating in on-call shifts, and mapping shift hours to inform the roadmap and create a source of truth for SOPs and training records. Defining 'done' for automation workflows, establishing metrics for toil reduction, and creating customer-facing dashboards for site SLOs and deployment cycle times are also critical. The role includes running the queue for four Decision Engineers, owning prioritization, and prototyping ideas with AI tools. The ultimate goal is to improve efficiency, measured in employee-hours per gigawatt.

Requirements

  • Shipped technical products for operations or infrastructure users, often after starting as an engineer or operator.
  • Owned a product measured on operational numbers: uptime, MTTR, machines per technician, audit findings closed, and not on launches.
  • Been the first product person in a function and set the roadmap, the metrics, and the working cadence from nothing.
  • Prototype with AI tools daily and ship what works instead of waiting on an engineering queue.
  • Credible with production engineers and SREs on the rotation and with leadership in the same hour, and have gotten both to change how they work.
  • Owned prioritization for a team of engineers: decided what shipped next and defended why.
  • Product and design taste shows in what you've shipped: interfaces the people on the rotation call obvious, workflows that survive a 2am page.

Nice To Haves

  • Datacenter or GPU fleet operations.
  • CMMS, DCIM, or asset management rollouts.
  • SRE and observability tooling.
  • Hardware qualification or burn-in.
  • Compliance or audit readiness.
  • Time on an on-call rotation.

Responsibilities

  • Own the automation roadmap for compute production as its first product manager.
  • Decide which parts of keeping GPU fleets worth billions healthy become software next, across fleet health, repair and RMA, hardware qualification, on-call, facility maintenance, and the asset model.
  • Defend the order with numbers: machines per operator, time to return to service, pages per failure mode.
  • Land the systems already in flight: a maintenance system for lockout tagout and work orders, an asset register, and retiring the legacy datacenter inventory.
  • Sequence those three and call the cutover dates within the first quarter.
  • Live on the floor and the rotation: embed with production engineers and facility operators, sit the on-call shift, map where shift hours actually go.
  • Turn that map into the roadmap everyone can cite, including the SOP source of truth and training records.
  • Define done for every workflow: the number each automation must clear before it counts, published where the team sees its own toil falling.
  • Create site SLO and deployment cycle time dashboards the customer reads.
  • Run the queue for the four Decision Engineers hired into this function.
  • Own what ships next and why.
  • Prototype own ideas with AI tools (Claude Code, Cursor, LLM APIs, MCP) and put working software into a shift lead's hands instead of writing requirements.
  • Answer for the outcome, employee-hours per gigawatt, with the Decision Engineers.

Benefits

  • Commitment to pay equity and transparency.
  • Equal Employment Opportunity Employer status.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service