About The Position

Propio Language Services is a leading provider of multilingual interpretation, translation, and localization services, operating at a significant scale across various industries. We are seeking a Senior Machine Learning Engineer, Speech & LLM Training Data to play a crucial role in transforming large volumes of multilingual conversational audio into high-quality training and evaluation datasets. This is a hands-on position responsible for audio processing, dataset curation, annotation, quality assurance workflows, model training, and evaluation for our multilingual speech, translation, and conversational AI systems.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or a related field, or equivalent practical experience.
  • 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows.
  • Strong hands-on experience with Python, SQL, Linux, Git, and Docker.
  • Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks.
  • Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools.
  • Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis.
  • Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines.
  • Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS.
  • Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling.
  • Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS.
  • Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.

Nice To Haves

  • Experience with multilingual telephony, healthcare, interpretation, or call-center audio.
  • Experience with tools such as Silero VAD, pyannote, WhisperX, NeMo, Kaldi, or equivalent speech technologies.
  • Experience with distributed processing or training using Ray, PySpark, or similar frameworks.
  • Experience with HIPAA, PHI/PII redaction, and secure data governance.
  • Experience with low-resource languages, accents, dialects, and code-switching.
  • Experience with synthetic data, active learning, weak supervision, or LLM-as-judge evaluation.

Responsibilities

  • Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI.
  • Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering.
  • Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage.
  • Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows.
  • Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements.
  • Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation.
  • Productionize secure, traceable, and reproducible data and ML workflows on AWS.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service