Cantina Labs is a social AI company developing advanced real-time models for expression, personality, and realism to bring characters to life and transform storytelling, connection, and creation. Cantina, our flagship social AI platform, is the beginning of our mission to shape human creativity and social interactions through AI. We are seeking a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end, with a focus on joint audio-video modeling. This role involves owning the audio side of multimodal generation, including representations, generative backbones, and conditioning/alignment for characters to speak, sing, and emote in sync with visuals. Responsibilities include voice cloning, multi-speaker conditioning, cinematic dialogue with music and sound design, and adjacent speech tasks. The role drives the model-data-evaluation flywheel, collaborating with research, video, data, and infra teams to ship fast, reliable, and cost-aware models at the intersection of research and engineering, contributing to safe, steerable, and trustworthy AI systems.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed