You'll build the distributed systems that train Luma's large-scale multimodal models across thousands of GPUs, so researchers can focus on innovation on top of reliable, efficient, scalable infrastructure. This is hard PyTorch, CUDA, and distributed-systems work — advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models at scale. If you haven't worked at the level of FSDP and multi-node training, this is the wrong depth.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed