Research Engineers on ML Infra make our models train and run performantly, reliably, and at scale by owning the distributed systems that sit underneath every model our researchers ship. As a ML Infra Research Engineer, you will: Architect, debug, and optimize the distributed ML training stack across the model, layer, and kernel levels — eliminating runtime and reliability bottlenecks on large GPU clusters. Profile end-to-end training runs to find bottlenecks across compute, communication, and storage, and build tooling to monitor throughput, utilization, and uptime across clusters. Optimize ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels. Work closely with Research Scientists to ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs. Own reliability of the training stack: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed