Cerebras Systems builds the world's largest AI chip, the Wafer-Scale Engine (WSE), which is 56 times larger than GPUs. This architecture enables industry-leading training and inference speeds, transforming AI applications and unlocking real-time iteration. Cerebras collaborates with leading AI organizations, including a significant partnership with OpenAI. The company is seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate its advanced machine learning compute clusters. This role offers the opportunity to work with the WSE and its supporting systems, ensuring the health, performance, and availability of the infrastructure, maximizing compute capacity, and supporting AI initiatives. The position requires a strong understanding of Linux-based systems, containerization, and experience with monitoring and troubleshooting distributed systems. The ideal candidate is a proactive problem-solver with expertise in large-scale compute infrastructure, dependability, and a commitment to customer success.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed