Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This involves rethinking every layer of the stack, from acquiring power and designing data centers to operating them with teams spanning hardware and software. Speed and scale are key differentiators. The company is looking for individuals who care deeply about this problem space and are motivated to contribute to building this infrastructure. Fluidstack operates with principles of extreme ownership, full autonomy, velocity, first principles thinking, and a passion for the problem space. The Production Engineering Team is working on critical problems such as building a repair pipeline for a large fleet of GPUs, qualifying new GPU generations within tight deadlines, migrating live compute at construction speed, and developing the observability and orchestration layer for hyperscale AI compute. This role is focused on owning the compute fleet health end-to-end, building the necessary metrics pipelines, alerting, and health views. It involves transforming deployment and repair into automated pipelines, designing and expanding the GPU qualification platform, and owning Redfish and BMC tooling for firmware-level telemetry and low-level access. The ultimate goal is to ensure the end-to-end reliability, scalability, and operation of the compute fleet at scale through aggressive automation, tooling, and incident discipline.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed