About The Position

MWDN is seeking an Applied Machine Learning Engineer specializing in audio and generative AI for an AI-native music label. This company uses artificial intelligence to power its entire content-production process, including artist-consistent audio generation, vocal cloning, custom text-to-speech, and generative lyrics. They have developed an end-to-end music production pipeline. This role is a practical engineering position focused on turning sophisticated AI models into release-ready music products. The engineer will take ownership of the workflow, which includes audio data preparation, model training, vocal conditioning, inference, quality evaluation, and deployment.

Requirements

  • 3+ years of experience training and deploying deep learning models in production, preferably within audio or NLP.
  • Hands-on experience in at least two of the following areas: Voice cloning or Singing Voice Conversion (SVC), Text-to-Speech model fine-tuning, LLM fine-tuning using techniques such as LoRA, DPO, and RAG.
  • Strong proficiency in PyTorch.
  • Practical experience with GPU cloud infrastructure, such as RunPod or AWS, as well as Docker and serverless inference.
  • A rigorous, evidence-based approach to model evaluation, including benchmarks, ablation studies, and blind testing.
  • Native or strong professional proficiency in Spanish, required for lyrics evaluation and voice quality assurance.

Nice To Haves

  • A music background or experience with music-production tools and workflows, including stems, MIDI, and DAWs.
  • Experience with singing-voice synthesis solutions such as ACE Studio, ACE-Step, RVC, so-vits-svc, or similar technologies.
  • Previous experience working with licensed celebrity or artist voices and consent-based voice AI.

Responsibilities

  • Fine-tune and improve an existing singing-voice conversion and cloning model, increasing its current 70–75% fidelity baseline to production-level quality of 90% or higher in blind listening tests.
  • Reproduce expressive vocal characteristics beyond basic timbre, including vibrato, falsetto, dynamics, and both spoken and sung delivery.
  • Evaluate different improvement strategies for singing voice models, including expanding the training dataset, testing alternative base models, and assessing enterprise APIs.
  • Take ownership of an existing validated LoRA fine-tune based on Chatterbox Multilingual for a Spanish-speaking voice.
  • Ensure accurate reproduction of a Mexican accent and correct pronunciation of Spanish sounds and letters.
  • Containerize the Text-to-Speech model and deploy it as a serverless inference endpoint using RunPod or a similar platform.
  • Integrate the Text-to-Speech endpoint with the existing web platform through an API.
  • Train a second version of the Text-to-Speech model using clean studio recordings to eliminate remaining output instability.
  • Develop a proprietary Spanish-language lyrics-generation model, with experience in regional Mexican genres considered a strong advantage.
  • Build the lyrics-generation solution using an open-weight LLM, LoRA fine-tuning, preference optimization through DPO, and RAG based on a curated content corpus.
  • Establish clear evaluation criteria and organize quality assessment for the lyrics model with native Spanish speakers.
  • Integrate the completed lyrics model into the existing frontend.

Benefits

  • Flexible working hours
  • 29 days of PTO (18 working days per year plus all national holidays)
  • 10 paid recovery days
  • Full financial and legal support for independent contractors
  • Free English classes, with native speakers or Ukrainian teachers
  • Dedicated HR support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service