Join a fast-moving AI infrastructure team working on the cutting edge of large-scale ML workloads. This role is ideal for engineers who enjoy solving deep technical challenges in distributed training, multi-GPU systems, and scalable AI inference infrastructure. You will work directly with AI-focused clients, helping them get the most out of modern GPUs (H100, B200, etc.) and ML frameworks such as PyTorch (and JAX in some environments). Work alongside senior AI and infrastructure engineers building large-scale GPU platforms. As part of the customer solutions team, you will design and validate production-grade distributed training (primary) and large-scale inference architectures on large GPU clusters, typically tens to thousands of GPUs. You will work hands-on with customers to debug, optimize, and scale ML workloads across multi-node GPU environments. Act as a technical authority on GPU performance, networking, and schedulers, making trade-offs at scale and translating customer needs into concrete platform requirements. Collaborate closely with engineering, product, and R&D to influence roadmap decisions based on real-world ML workloads. This is a hands-on, technical role; you are expected to work directly in customer environments, not only advise at a high level.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed