Our client is building a new generation of highly efficient AI models designed to dramatically improve the speed and economics of large language model inference. The company has pioneered diffusion-based language models that generate responses in parallel rather than relying exclusively on traditional sequential token generation. This approach enables significantly faster and more efficient AI inference while maintaining competitive model quality. The company launched one of the first commercially available diffusion-based language models in early 2025 and is now deploying large-scale AI models with Fortune 500 organizations. The team is small, highly technical, and research-driven, with engineers working directly alongside world-class researchers and founders. The organization places a strong emphasis on technical depth, experimentation, performance optimization, and production-scale AI infrastructure. As a Machine Learning Systems Engineer, you'll work on the infrastructure that enables large-scale model training and inference, contributing directly to systems that make advanced AI models faster, more efficient, and more reliable. This is an opportunity to join an elite AI team where you can work at the intersection of machine learning, distributed systems, GPU infrastructure, and high-performance model serving.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior