Research Scientist, Video Foundation Models

Cantina•San Francisco, CA
•$200,000 - $320,000•Onsite

About The Position

We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation. In this role, you will work on foundational research and large-scale model development across the full model lifecycle, including architecture, data, training, evaluation, post-training, training systems, and efficient inference. You will have the opportunity to shape both the technical direction and the team from an early stage.

Requirements

  • Strong research and engineering experience in generative modeling, including areas such as diffusion models, flow matching, DiTs, video generation, multimodal models, world models, or related fields.
  • Hands-on experience training and evaluating large-scale image, video, or unified multimodal models using modern deep learning frameworks and distributed training systems.
  • A strong track record of developing impactful models or systems, demonstrated through research publications, open-source contributions, production impact, or other significant technical work.
  • Ability to independently own ambiguous research problems, move effectively from ideas to experiments, and work well in a highly collaborative environment.
  • Specialized depth in one or more areas across the foundation model lifecycle, such as model architecture, data curation, controllable generation, multimodal understanding and conditioning, post-training and reward modeling, model acceleration, inference systems, or deployment.

Responsibilities

  • Research, develop, and scale native video and multimodal foundation models, from early prototypes through large-scale pre-training, continued training, and post-training.
  • Explore new model architectures, training objectives, and conditioning mechanisms for video generation, reference- and memory-based generation, multimodal interaction, and joint audio-video generation.
  • Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines for high-quality and controllable generation.
  • Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency.
  • Collaborate closely with researchers, engineers, and product teams to help shape the technical roadmap and, where appropriate, translate model advances into real-world capabilities.
  • Contribute to research publications and open-source releases when appropriate.

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including: 15 PTO days, 10 sick days, 15 company holidays, 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service