Product Engineer, Compute Operations

Fluidstack•New York, NY
•$224,000 - $300,000•Onsite

About The Position

Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else by rethinking every layer of the stack. This involves acquiring power, designing and building data centers, and operating them with integrated hardware and software teams. The company prioritizes speed and scale as key differentiators and seeks individuals who are deeply committed to this problem space. The company operates with a philosophy of full autonomy, insane urgency, reasoning from first principles, and a passion for tackling the frontier of AI. The Decision Team is focused on automating the delivery of gigawatts by turning processes into software, forward-deploying alongside experts to integrate their judgment into systems, and ensuring every supercomputer is delivered faster than the last by leveraging a shared knowledge graph for continuous learning and improvement.

Requirements

  • Shipped production code in Go, Python, or TypeScript, and can pick up whatever language the problem demands.
  • Built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks.
  • Works daily with AI coding tools like Claude Code and Cursor, and gets agents doing useful work autonomously alongside them.
  • Identifies problems, designs the solution, and ships it without waiting for direction or approval.
  • Moved fast under deadline while leaving foundations that other engineers extended after you moved on.
  • Sat the on-call rotation or worked beside the people who do, and turned operational pain into systems that made the pager quieter.
  • Product taste shows in what you've shipped: interfaces the engineers on the rotation call obvious, and workflows that match how the work actually happens.

Nice To Haves

  • Production engineering or SRE on large GPU fleets.
  • Hardware qualification or burn-in frameworks.
  • BMC, Redfish, or IPMI tooling.
  • CMMS, DCIM, or asset management systems.
  • BMS/EPMS or SCADA.
  • Prometheus and Grafana.

Responsibilities

  • Build the fleet health system: real-time telemetry and tiered healthchecks on every machine across Kubernetes and bare metal, rolled into one API the whole company trusts to answer "is this machine healthy," with alarms correlated into incidents that reach on-call with a drafted probable cause.
  • Turn repair and RMA into generated work: one tracked flow from failure detection through triage, parts, vendor return, and return to service, where failure thresholds route machines to repair automatically, each production engineer's shift todo list is generated for them, and time to return to service is a number the system reports.
  • Ship hardware qualification as software: burn-in, performance baselining, and new hardware validation composed into rack-level workflows, so bringing thousands of accelerators online is a repeatable run and every machine enters production with its acceptance evidence attached in the graph.
  • Run the facility on the same system as the fleet: the maintenance system for lockout tagout and work orders is live at one site and rolls out to two more, every asset register loads before the first external audit this fall, and the legacy datacenter inventory retires before the next building energizes. You own the asset model, the migration, and the day the old tools switch off.
  • Turn every runbook into a checked procedure: SOPs, training records, and technician qualifications become structured data the customer can audit, and site SLOs, deployment cycle time, and labor ramp report themselves on the dashboards a hyperscaler customer asked for. You work forward-deployed beside production engineers and facility operators, on site and on the rotation, and build what they use the next shift.

Benefits

  • Competitive total compensation package (cash + equity)
  • Health, dental, and vision insurance
  • Retirement plan
  • Generous PTO policy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service