Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This involves rethinking every layer of the stack, from acquiring power to designing, building, and operating data centers. The company operates with a focus on extreme ownership, velocity, first principles, and a deep passion for the problem space. The Data Center Operations Team is tasked with operating at the scale of a nation, managing a fleet that will draw more power than some countries, and bringing sites online rapidly while maintaining flawless operation of live systems. The team is responsible for writing the operational playbook for this unprecedented speed and scale. This role specifically leads the compute production engineering team responsible for keeping tens of thousands of GPUs serving customers. The lead will own fleet availability for compute, defining SLOs, building tooling, and improving metrics. Key responsibilities include building automation for the entire node lifecycle (provisioning, health checks, remediation, return to service) without human intervention, and setting up an effective on-call and escalation model that maintains sharp response times without causing burnout.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed