The ML Training Performance Engineer will own the efficiency and scalability of training models on GPU and cloud infrastructure. You will turn available compute into faster, more capable experiments as our model training scales in complexity and compute requirements. As a hands-on individual contributor on a small research team, you will work across the training stack, from Python and model execution to GPU kernels, distributed communication, and runtime environments. Partnering with model researchers, data engineers, and infrastructure and IT teams, you will identify and implement performance improvements while preserving numerical correctness and scientific intent. You will have significant ownership over how we measure, optimize, and scale training performance, establishing the baselines, tooling, and technical approaches that will support our AI/ML work as it grows.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Associate degree