Network Production Engineering Lead

FluidstackSan Francsisco, CA
$176,000 - $221,000

About The Position

Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This role is within the Data Center Operations Team, which operates at the scale of a nation, not a building, and manages live sites while construction continues. The team is responsible for writing the playbook for operating at unprecedented speed and scale. This specific role involves leading the network production engineering team responsible for maintaining the health of fabrics for 100k+ accelerator clusters. The lead will own network availability and performance SLOs, including link health, congestion, and failure response. Key responsibilities include building automation for fabric operations such as telemetry, anomaly detection, and automated drain and repair, as well as establishing the operating model between design engineering and site operators for clean escalation flows.

Requirements

  • Led network operations or production engineering for very large fabrics.
  • Automated network remediation at scale and trusted it enough to let it run.
  • Read fabric telemetry and find the sick link before the training job does.
  • Run on-call programs teams didn't hate.

Nice To Haves

  • AI or HPC fabrics.
  • InfiniBand and RoCE.
  • Network telemetry stacks.
  • Vendor TAC escalation management.

Responsibilities

  • Lead the network production engineering team keeping fabrics for 100k+ accelerator clusters healthy.
  • Own network availability and performance SLOs: link health, congestion, and failure response.
  • Build automation for fabric operations: telemetry, anomaly detection, automated drain and repair.
  • Set the operating model between design engineering and site operators so escalations flow clean in both directions.

Benefits

  • Pay equity and transparency.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service