Fluidstack is building civilization-scale infrastructure for AI, aiming to deliver 10 to 100s of GWs of compute faster than anyone else. This involves rethinking every layer of the stack, from acquiring power to designing, building, and operating data centers. The company operates with a philosophy of extreme ownership, full autonomy, velocity, first principles thinking, and a deep passion for the problem space. The Data Center Operations Team is responsible for operating at the scale of a nation, managing sites that come online in pieces while ensuring live operations remain flawless. They are tasked with writing the playbook for operating at this unprecedented speed and scale, as no prior operations organization has done so. The Customer Reliability Engineer role focuses on owning the reliability for named customer workloads, including their clusters, SLAs, and escalations. This involves debugging across the full stack (hardware, fabric, scheduler) when training runs degrade, managing customer-facing incident communication with technical depth and transparency, and transforming recurring customer pain points into engineering fixes with the production teams.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed