Applied Intuition is seeking a performance engineer specializing in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role focuses on distributed training runs spanning many nodes and high-throughput batch inference for processing large datasets. The primary optimization targets are throughput, cluster goodput, and cost per unit of data processed, aiming to reduce inefficiencies in GPU-hours and processing time. The engineer will be responsible for profiling across the stack, identifying performance bottlenecks, and implementing solutions to bridge the gap between theoretical accelerator capabilities and actual workload performance. This involves working at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, collaborating with various teams to improve training time-to-result and offline processing costs. Engineers at Applied Intuition are encouraged to take ownership of technical and product decisions, interact with users, and contribute to a dynamic team culture.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed