This role is critical for bridging the gap between frontier research and production reality in a fast-paced generative AI company. The primary goal is to remove the bottleneck in productionizing research checkpoints, ensuring faster API deployment, optimized inference speed, and robust API performance under load. The engineer will own the development and maintenance of high-performance inference services and APIs, optimize GPU infrastructure for latency and throughput, build scalable serving architectures, and improve system reliability and observability. This position involves close collaboration with researchers to rapidly move from idea to live endpoints and prototype demos. The role requires a strong understanding of backend systems, GPU performance, and production ML serving. The ideal candidate will have experience building and operating systems at scale, understanding the difference between research prototypes and production systems, and making informed tradeoffs regarding performance, reliability, and cost. Comfort in a fast-moving, research-adjacent environment and a strong sense of ownership are essential.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed