Our research team moves fast. Models improve weekly. New capabilities emerge constantly. What slows us down is not model quality—it’s productionization. Without this role: Research checkpoints sit longer before becoming usable APIs Inference is slower than it needs to be APIs struggle under load Demos don’t reflect the true potential of our models This role removes the bottleneck between frontier research and production reality. Once hired, researchers ship faster, demos launch faster, and customers experience models at their best. You will own the bridge between research breakthroughs and production systems. Turn research checkpoints into production-ready inference services Design and maintain high-performance APIs serving millions of requests Optimize inference latency and throughput across GPU infrastructure Build scalable serving architectures that handle unpredictable traffic Improve reliability, monitoring, and observability across model-serving systems Prototype and ship demos that showcase new capabilities in days, not weeks Collaborate closely with researchers to move from idea to live endpoint rapidly This role spans backend systems, GPU performance, and production ML serving.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed