Research, Vision Expertise

Thinking Machines LabSan Francisco, CA
$350,000 - $475,000Onsite

About The Position

Thinking Machines builds multimodal-first AI that extends human will and judgment. We are looking for new team members to advance the science of visual perception and multimodal learning. We focus on how vision and language interact at scale, designing architectures that fuse pixels and text, building datasets and evaluation methods for real-world comprehension, and developing representations that ground abstract concepts in the physical world. The goal is to create multimodal systems that integrate seamlessly into real-world environments. This role involves working at the intersection of visual understanding, multimodal reasoning, and large-scale model training, contributing to the development of architectures, data, and evaluation tools that enable AI to see, understand, and collaborate. The ideal candidate is curious about multimodal interfaces, experienced in running large-scale experiments, and comfortable contributing to complex engineering systems. While expertise in multimodality is sought, Thinking Machines Lab operates as a unified team, expecting new hires to work across modalities collaboratively. This role combines fundamental research and practical engineering, requiring high-performance code writing and the ability to read technical reports. It suits individuals who enjoy both deep theoretical exploration and hands-on experimentation, aiming to shape the foundations of AI learning. This is an 'evergreen role' kept open continuously to express interest in this research area, with applications reviewed on an ongoing basis for emerging opportunities.

Requirements

  • Ability to design, run, and analyze experiments thoughtfully, with demonstrated research judgment and empirical rigor.
  • Understanding of machine learning fundamentals, large-scale training, and distributed compute environments.
  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
  • Comfortable with debugging distributed training and writing code that scales.
  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Clarity in communication, an ability to explain complex technical concepts in writing.

Nice To Haves

  • Research or engineering contributions in visual reasoning, spatial understanding, or multimodal architecture design.
  • Experience developing evaluation frameworks for multimodal tasks.
  • Publications or open-source contributions in vision-language modeling, video understanding, or multimodal AI.
  • A strong grasp of probability, statistics, and ML fundamentals. Ability to distinguish between real effects, noise, and bugs in experimental data.
  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.

Responsibilities

  • Own research projects on training and performance analysis of multimodal AI models.
  • Curate and build large-scale datasets and evaluation benchmarks to advance vision capabilities.
  • Work with data infrastructure engineers, pretraining researchers and engineers, and the product team to create frontier multimodal models and the products that leverage them.
  • Publish and present research that moves the entire community forward.
  • Share code, datasets, and insights that accelerate progress across industry and academia.

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service