Software Engineer, Compute (GPU)

FluidstackAustin, TX
$208,000 - $269,000

About The Position

Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This involves rethinking every layer of the stack, from acquiring power and designing data centers to operating them with teams spanning hardware and software. Speed and scale are key differentiators. The company is looking for individuals who care deeply about this problem space and are motivated to contribute to building this infrastructure. Fluidstack operates with principles of extreme ownership, full autonomy, velocity, first principles thinking, and a passion for the problem space. The Production Engineering Team is working on critical problems such as building a repair pipeline for a large fleet of GPUs, qualifying new GPU generations within tight deadlines, migrating live compute at construction speed, and developing the observability and orchestration layer for hyperscale AI compute. This role is focused on owning the compute fleet health end-to-end, building the necessary metrics pipelines, alerting, and health views. It involves transforming deployment and repair into automated pipelines, designing and expanding the GPU qualification platform, and owning Redfish and BMC tooling for firmware-level telemetry and low-level access. The ultimate goal is to ensure the end-to-end reliability, scalability, and operation of the compute fleet at scale through aggressive automation, tooling, and incident discipline.

Requirements

  • Treat toil as a bug; manual steps in repair workflows are a backlog item.
  • Have an instinct for hardware, comfortable reasoning about failure modes at the firmware and silicon level.
  • Move toward ambiguity, build maps, and explain them.
  • Learn at a steep slope, reaching competence in unfamiliar domains quickly.
  • Carry a pager without flinching, run incidents, write postmortems, and fix systemic causes.
  • Be fluent with AI tooling (LLM APIs, MCP servers, agentic frameworks) and drive AI coding tools daily.
  • Have shipped production automation that other teams depend on.
  • Be comfortable in any language using AI coding tools.

Nice To Haves

  • Hardware lifecycle management and RMA automation.
  • BMC/Redfish or IPMI tooling.
  • GPU qualification or burn-in frameworks.
  • Workflow and orchestration engines (Temporal, Cadence).
  • Metrics and alerting pipelines (Prometheus, Grafana).
  • Experience with Go or Python.

Responsibilities

  • Own compute fleet health end to end, building metrics pipelines, alerting, and a unified health view for all GPUs in production.
  • Turn deployment/repair into a pipeline, building and owning automation for failure detection, triage, parts management, and return to service.
  • Design and expand the GPU qualification platform, including burn-in, performance baselining, and NPI execution for new GPU generations.
  • Own Redfish and BMC tooling for firmware-level telemetry, log collection, and the low-level access layer.
  • Own end-to-end reliability, scalability, and operation of the compute fleet at-scale.
  • Drive aggressive automation, tooling, and incident discipline to support the growth of one of the largest GPU fleets in the world.

Benefits

  • Competitive total compensation package (salary + equity).
  • Retirement or pension plan, in line with local norms.
  • Health, dental, and vision insurance.
  • Generous PTO policy, in line with local norms.
  • Equity in the form of stock options.
  • Commitment to pay equity and transparency.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service