This role focuses on making Luma's multimodal models fast by profiling and optimizing GPU, CPU, and accelerator code. The goal is to ensure efficient training and scalable deployment without compromising quality. The position involves writing kernels and operations to maximize hardware utilization. This is a deep performance-focused role requiring expertise in fused kernels, tensor cores, Triton, CUDA, and distributed multi-node deployment. A strong background in GPU optimization and a deep understanding of transformer internals are essential. Proficiency in CUDA, Triton, and profilers is a must.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed