Sr. Research Engineer

AdobeSan Francisco, CA

About The Position

Adobe's Sound Design AI group (SODA) is seeking a motivated Data/ML engineer to advance audio Generative AI. This role is part of the team responsible for Firefly Generate Sound Effects and other AI models integrated into Adobe products. The team is small, collaborative, and efficient, looking for highly motivated individuals with a passion for audio and video data. The position involves productizing cutting-edge research into tools for Adobe's creative users. Specifically, this role is a data lead responsible for the end-to-end training data for generative audio models. These models are trained on large-scale audio and video corpora, and the quality, balance, and integrity of this data are crucial for model performance. The individual in this role will be the sole owner of the data content, the rationale behind its selection, and the methods for ensuring its quality.

Requirements

  • Pipeline-building experience at scale for audio and/or video data with reproducibility and versioning.
  • Deep audio domain knowledge, including common datasets, quality metrics and their failure modes, codecs, normalization, and filtering.
  • Strong "ears" and critical listening ability.
  • Generative-ML research experience to make data decisions, design and run training experiments independently.
  • Strong command of evaluation in audio and video modeling.
  • Data acquisition and licensing experience, including writing briefs and communicating across stakeholders.
  • Excellent communication skills.

Nice To Haves

  • Video knowledge a plus.

Responsibilities

  • Build, deploy, and monitor large-scale data pipelines for ingesting, filtering, and preprocessing audio and video data at scale, ensuring reproducibility and versioning.
  • Run in-house models for inference over millions of audio/video files and develop data exploration tools to support this process.
  • Curate and select training data using both quantitative metrics and qualitative judgment.
  • Maintain a comprehensive understanding of the corpus, including its distribution, balance, diversity, coverage, and redundancy.
  • Train or finetune models on pipeline outputs, evaluate their behavior, and use these findings to guide new data experiments and ablations.
  • Scale training processes and optimize throughput.
  • Drive data licensing and acquisition by identifying gaps and opportunities, defining briefs, and collaborating with vendors and producers to license and curate new datasets.
  • Own the production, validation, and improvement of training labels and annotations, including the implementation of model-assisted labeling at scale.

Benefits

  • Comprehensive benefits programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service