Research Scientist, Video & Multimodal

Innodata Inc.
•$160,000 - $185,000

About The Position

Innodata is a global data engineering company focused on enabling the responsible advancement of artificial intelligence by providing essential data, evaluation frameworks, and human expertise. The company specializes in solutions, platforms, and services for Generative AI and AI builders. This role focuses on the challenging area of video and multimodal models, where temporal reasoning, long-form understanding, and the integration of audio, video, and text are critical but often underdeveloped. The Research Scientist will be instrumental in closing the gap in data design and evaluation for these models, working with customers and frontier labs to advance video and multimodal AI.

Requirements

  • Approximately 5+ years of hands-on industry experience in video understanding or multimodal ML. Practical experience is weighted heavily over formal credentials; a PhD with a compelling, current research agenda can offset lower experience.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required.
  • Experience training and evaluating video or multimodal models yourself, with strong PyTorch fundamentals.
  • Fluency in video processing formats and tooling, including ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models with the modern toolchain (HuggingFace transformers, PEFT, efficient inference).
  • Experience with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • A dataset and benchmark-oriented mindset: experience building evaluation sets, calibrating difficulty, and understanding what makes video data effective for specific objectives.
  • A recognized track record in the field, demonstrated through first-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR.
  • Ability to work directly with research scientists at customer and partner organizations.
  • Ability to clearly explain data and modeling decisions to both expert and non-expert audiences.
  • A rigorous, reproducible approach to experiments and documentation.

Nice To Haves

  • An advanced degree (MS or PhD) in a relevant field is preferred.
  • Interest or hands-on experience in responsible-AI evaluation and red-teaming, including safety and robustness testing for video and multimodal systems.

Responsibilities

  • Define how Innodata designs, structures, and evaluates video data for video and multimodal models, and validate these choices experimentally.
  • Translate requirements of video and multimodal models (e.g., video understanding, temporal grounding, video generation) into concrete data specifications, including modalities, annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology for video understanding, addressing temporal grounding accuracy, long-context reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluations, clarifying when model-based scoring is trustworthy versus when human judgment is needed.
  • Build evaluation methodology for video generation, focusing on fidelity, temporal coherence, and physical plausibility, including generative video used as a world model, in an area where automatic metrics are weak and human judgment is paramount.
  • Determine how existing and incoming video data should be structured, enriched, and sampled to maximize model value, including from complex, domain-specific footage.
  • Run experiments to demonstrate the impact of data decisions, including fine-tuning and evaluating models on Innodata data with ablations to link specific data choices to measurable improvements.
  • Design adversarial and stumping evaluations to identify weaknesses in video and multimodal systems and use these failures to improve data collection and design.
  • Publish research findings in the form of benchmarks, methodologies, and papers to advance the field and build trust with partners.
  • Collaborate with annotation teams, subject-matter experts, and the synthetic-data pipeline to translate specifications into operational collection and labeling plans.

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service