This role involves designing, deploying, and maintaining large distributed ML training and inference clusters. The position requires developing efficient, scalable end-to-end pipelines for managing petabyte-scale datasets and model training throughout the entire ML lifecycle. Responsibilities also include researching and testing various training approaches, analyzing and debugging low-level GPU operations for performance optimization, and staying current with research to introduce new ideas.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed