Thinking Machines is seeking an infrastructure research engineer to design, optimize, and maintain the compute foundations for large-scale language model training. This role involves developing high-performance ML kernels (e.g., CUDA, CuTe, Triton), enabling efficient low-precision arithmetic, and enhancing the distributed compute stack for training large models. The ideal candidate enjoys working closely with hardware and across research boundaries, collaborating with researchers and systems architects to bridge algorithmic design with hardware efficiency. Responsibilities include prototyping new kernel implementations, profiling performance across hardware generations, and defining numerical and parallelism strategies for next-generation AI systems. This is an evergreen role, meaning applications are continuously reviewed for current and future opportunities.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Associate degree