Software Engineer, API Multimodal

OpenAISan Francisco, CA
$293,000 - $385,000

About The Position

API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required.

Requirements

  • 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles.
  • A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed systems.
  • Strong software engineering and systems fundamentals, with experience leading technically complex projects from ambiguous ideas to production.
  • Proficiency in one or more general-purpose backend languages, such as Python, Go, Rust, or TypeScript.
  • Experience building reliable, scalable systems and reasoning about distributed architecture, concurrency, latency, observability, and operational tradeoffs.
  • Product judgment and developer empathy, including an ability to turn complex model or infrastructure capabilities into clear, intuitive APIs.
  • Clear communication and a collaborative approach to working with researchers, product managers, designers, infrastructure engineers, and customers.
  • A strong sense of ownership, comfort with ambiguity, and a bias toward learning directly from users while continuously improving engineering quality.

Nice To Haves

  • Experience with real-time streaming, audio processing, speech systems, image generation, computer vision, or multimodal AI applications is helpful, but not required.

Responsibilities

  • Design, build, and ship developer-facing APIs and backend services that serve frontier models.
  • Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale.
  • Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers.
  • Own the availability, latency, scalability, and cost efficiency of the services you build.
  • Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service