Technical Lead, Multimodal Research

Eventual•San Francisco, CA
•Onsite

About The Position

Eventual is building the infrastructure for Physical AI, enabling teams to find and utilize specific situations across vast amounts of video, lidar, radar, and sensor data. Their open-source engine, Daft, is designed for multimodal AI, handling petabytes of data daily. The company aims to make indexing and understanding this data more efficient and cost-effective than traditional annotation methods. Eventual has raised $30M and its team comprises experienced professionals from leading tech companies. They are looking for individuals to join their small, powerful team, working 4 days a week in their SF Mission District office.

Requirements

  • 5+ years in applied computer vision or multimodal ML.
  • PhD or MS in computer science, electrical engineering, robotics, or applied mathematics with a computer vision or machine learning focus, or a comparable publication or production record.
  • Depth in modern vision and multimodal modeling (VLMs, VQA, embeddings, representation learning, detection, tracking, segmentation, retrieval) with a focus on deployable solutions.
  • Hands-on training and evaluation of models at scale on real video and sensor data.
  • Comfort across the research and engineering boundary, including PyTorch prototyping, inference performance, GPU utilization, throughput, and cost.
  • Background from a perception or multimodal team at a self-driving, robotics, or Physical AI company, a frontier research lab, or a visual-data company, ideally as the senior-most person on that problem.

Nice To Haves

  • Publications at CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR.
  • Experience building or fine-tuning VLMs or other multimodal foundation models.
  • Experience with long-form video, temporal reasoning, embeddings, retrieval, or content-aware indexing at scale.
  • Experience with multimodal sensor data beyond RGB (lidar, radar, depth, or simulation output).
  • Experience with evaluation frameworks, labeling taxonomies, large-scale annotation programs, or inference and training optimization across large GPU clusters.

Responsibilities

  • Own the execution of the technical vision for understanding video data.
  • Decide which models, representations, and evaluation methods to use.
  • Prove technical approaches in production at petabyte scale.
  • Stay hands-on with papers, models, and experiments.
  • Own modeling strategy across the platform, including model families, representations, and training approaches.
  • Take approaches from prototype into production inference at corpus scale, collaborating with data systems and storage teams.
  • Define the evaluation standard for models before they reach customers.
  • Own the cost curve for understanding through architectural decisions like distillation, cascades, routing, and quantization.
  • Translate customer research needs into scoped technical programs and set the technical direction for multimodal work.

Benefits

  • In-person, tight-knit team — 4 days/week in our SF Mission office.
  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team-building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment.
  • 401(k) plan with match.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service