Research Scientist, Speech & Audio

Innodata Inc.
•$160,000 - $185,000

About The Position

Innodata is a global data engineering company focused on the intersection of data and Artificial Intelligence (AI). Our mission is to enable the responsible advancement of AI by providing essential data, evaluation frameworks, and human expertise for building trustworthy AI systems at scale. We offer a range of solutions, platforms, and services for Generative AI/AI builders and adopters, leveraging our 36+ year legacy of delivering high-quality data and outstanding customer outcomes. This role focuses on the scientific aspects of speech and audio data for AI models. You will partner directly with customers and frontier labs working on ASR, text-to-speech, speech-to-speech, conversational voice, diarization, and audio-language models. Your work will involve critical judgment on benchmark design, transcription conventions, and the appropriate use of automated versus human evaluation metrics. You will also collaborate closely with our transcription and linguistics lead to ensure model learning is shaped by precise standards.

Requirements

  • Approximately 5+ years of hands-on industry experience in speech or audio ML; practical experience is weighted over formal credentials, though a PhD with a compelling research agenda can offset lower experience.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required.
  • Experience training and evaluating speech or audio models (ASR, TTS, speech-to-speech, speaker, or audio-language models) with strong PyTorch fundamentals.
  • Fluency in toolchains and metrics used in speech work, such as ESPnet, NeMo, SpeechBrain, Kaldi, HuggingFace, forced alignment, and WER/CER and advanced metrics.
  • Hands-on experience with multilingual, accented, dialectal, low-resource, or code-switched speech, and with synthetic or augmented audio (TTS pipelines, noise and room-response simulation).
  • A dataset-centric approach to thinking, including experience building evaluation sets, reasoning about coverage across conditions, and understanding what makes speech data effective for a given objective.
  • A recognized track record in the field, demonstrated through first-author publications or strong open-source contributions at venues like Interspeech, ICASSP, ASRU, SLT, or NeurIPS.
  • The ability to work directly with research scientists at customer organizations and frontier labs, and to clearly explain data and modeling decisions to both expert and non-expert audiences, supported by a rigorous, reproducible approach to experiments and documentation.

Nice To Haves

  • An advanced degree (MS or PhD) in a relevant field is preferred.
  • Interest or hands-on experience in responsible-AI evaluation and red-teaming, such as spoofing and voice-cloning robustness, or bias across accents and languages.

Responsibilities

  • Define how Innodata designs, structures, and evaluates audio data for speech and audio models, and validate these choices experimentally.
  • Translate the requirements of various speech and audio models (ASR, text-to-speech, speech-to-speech, conversational voice, speaker diarization and verification, audio-language models, streaming systems) into concrete data specifications, including modalities, transcription and annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodologies that go beyond word error rate, encompassing semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency, and determine when automated metrics are reliable.
  • Decide how existing and incoming audio should be structured, enriched, and sampled to achieve coverage that aligns with model objectives, considering languages, accents, acoustic conditions, speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech.
  • Partner with the transcription and linguistics lead to translate model objectives into transcription specifications and quantify the impact of transcription conventions and quality on ASR and speech-model results.
  • Collaborate with the audio solutions and engineering team to ensure collected audio is optimized for model objectives, specifying requirements for good data and evaluation, and overseeing program scoping and audio capture.
  • Run experiments to demonstrate the impact of data decisions, including fine-tuning and evaluating models on Innodata data with ablations to link specific data choices to measurable improvements.
  • Design adversarial and stumping evaluations (noisy, accented, adversarial audio) to identify weaknesses in speech systems and use these findings to improve data.
  • Publish research findings in the form of benchmarks, methodologies, and papers to advance the field and build trust with partners.
  • Work with annotation teams, subject-matter experts, and the synthetic- and augmented-audio pipeline to translate specifications into operational plans.

Benefits

  • Competitive salary range of $160,000 - $185,000 p/year, based on experience, skills, and qualifications.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service