Applied Research Scientist, AI Research

DescriptSan Francisco, CA
Hybrid

About The Position

Descript’s Research team builds the generative video and audio models, and the multimodal understanding systems, behind Descript's most distinctive features: Video Regenerate and lipsync, video translation, and zero-shot voice and roomtone cloning. This isn't research for its own sake — everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months. We're a small, senior-heavy team, and we're always looking for strong applied research scientists — generalists in deep learning and generative modeling, not narrow specialists.

Requirements

  • Proven ability to design and implement deep learning algorithms, demonstrated by publications, open-source work, or models you've shipped.
  • Strong programming skills and deep fluency in PyTorch and/or TensorFlow.
  • A track record of generating new ideas in machine learning — you produce more ideas than you can implement, and once an experiment setup is established, you can run and evaluate many of them quickly rather than being bottlenecked on infrastructure.
  • Strong experimental judgment: you test ideas fast, and you're honest with yourself and the team about which ones don't pan out.
  • A PhD or Master's in deep learning or a related field, or equivalent experience — we care about the track record more than the credential.
  • At least one of the following must be true: Lead or first author of an accepted publication in a top venue: NeurIPS, ICML, ICLR, ICASSP, ICCV, CVPR, Interspeech, SIGGRAPH, or similar.
  • Played a key role in shipping a production feature with deep learning as a core component.
  • More senior candidates (Senior and Staff) should also bring a track record of owning research direction rather than executing a plan handed to them, and experience mentoring or technically leading other researchers or engineers.

Nice To Haves

  • We don't require domain-specific expertise in computer vision or speech/audio specifically — our team spans both, and strong general deep learning ability transfers.
  • Depth in any of these is a strong signal: Generative modeling for video, audio, or images.
  • Vision-language models and multimodal understanding.
  • Speech and audio modeling, including compression, enhancement, and synthesis.
  • Building evaluation systems for generative or agentic outputs where metrics resist clean definitions.
  • Taking a research idea through to a shipped, production-facing feature.

Responsibilities

  • Generative media synthesis: building the models behind features like Video Regenerate, lipsync, and video translation — photorealistic video and audio synthesis as production research, not demos.
  • Zero-shot voice and roomtone cloning: realistic cloning from only a few minutes of reference audio.
  • Multimodal understanding: vision-language systems that let Descript's agentic editing features reason over visual and audio content, plus the evals that balance quality against cost and latency.
  • Computer-vision-heavy problems: digital human reconstruction, facial modeling, and related work in service of more natural editing tools.
  • New algorithms: media synthesis, speech enhancement, anomaly detection, and audio/video tagging.
  • Direction-setting: identifying and pursuing the next research direction that should become a Descript feature — not just a paper. (More senior candidates should expect to own this directly; more junior candidates will grow into it.)

Benefits

  • generous healthcare package
  • 401k matching program
  • catered lunches
  • flexible vacation time
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service