The Omni team at xAI creates magical AI experiences beyond text, enabling understanding and generation of content across various modalities, including image, video, and audio. As a multimodal engineer with a focus on audio, you will drive the model’s audio generation and understanding capabilities, particularly in the context of video and broader media generation. This includes advancing audio for video, music generation, general audio synthesis, speech processing, audio understanding, audio representation / tokenizer, and so on. You will contribute across data curation, modeling, inference serving, and product integration, working on both pretraining and post-training while collaborating with product teams to push the frontiers of model capability and end-to-end user experience.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed
Number of Employees
1,001-5,000 employees