Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This role is within the Data Center Operations Team, which operates at the scale of a nation, not a building, and manages live sites while construction continues. The team is responsible for writing the playbook for operating at unprecedented speed and scale. This specific role involves leading the network production engineering team responsible for maintaining the health of fabrics for 100k+ accelerator clusters. The lead will own network availability and performance SLOs, including link health, congestion, and failure response. Key responsibilities include building automation for fabric operations such as telemetry, anomaly detection, and automated drain and repair, as well as establishing the operating model between design engineering and site operators for clean escalation flows.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed