Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. The Core ML team develops novel machine learning algorithms that take advantage of the unique capabilities of the Cerebras Wafer-Scale Engine. Our work spans efficient LLM training and inference, parallel and diffusion-based generation, sparsity, scaling laws, and training dynamics. We are looking for an engineer to bridge the gap between promising research ideas and efficient execution on Cerebras systems. You will work across ML frameworks, compilers, runtimes, and low-level kernels to implement new algorithmic capabilities, diagnose performance bottlenecks, and turn research prototypes into robust, high-performance demonstrations. Depending on your background, your work may emphasize runtime capabilities such as token orchestration, scheduling, communication, and distributed execution; low-level kernel development for novel ML operations; or a combination of both.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level