This role is focused on designing and operating distributed training systems for large neural networks across GPU clusters. The engineer will optimize multi-node, multi-GPU execution, diagnose and resolve bottlenecks, and improve training stability and fault tolerance at scale. This position involves partnering with research and applied ML teams to productionize large-model training pipelines, ultimately building the foundation for large-scale AI.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed