Thinking Machines builds multimodal-first AI that extends human will and judgment. We are looking for new team members to advance the science of visual perception and multimodal learning. We focus on how vision and language interact at scale, designing architectures that fuse pixels and text, building datasets and evaluation methods for real-world comprehension, and developing representations that ground abstract concepts in the physical world. The goal is to create multimodal systems that integrate seamlessly into real-world environments. This role involves working at the intersection of visual understanding, multimodal reasoning, and large-scale model training, contributing to the development of architectures, data, and evaluation tools that enable AI to see, understand, and collaborate. The ideal candidate is curious about multimodal interfaces, experienced in running large-scale experiments, and comfortable contributing to complex engineering systems. While expertise in multimodality is sought, Thinking Machines Lab operates as a unified team, expecting new hires to work across modalities collaboratively. This role combines fundamental research and practical engineering, requiring high-performance code writing and the ability to read technical reports. It suits individuals who enjoy both deep theoretical exploration and hands-on experimentation, aiming to shape the foundations of AI learning. This is an 'evergreen role' kept open continuously to express interest in this research area, with applications reviewed on an ongoing basis for emerging opportunities.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior