DoorDash’s GenAI Platform team, within Machine Learning Platform, is responsible for building the shared infrastructure that enables teams across DoorDash, Wolt, and Deliveroo to safely deploy Generative AI-powered products, agents, automation, and personalization into production. The team's mission is to accelerate the business impact derived from GenAI. A key aspect of this work involves self-hosting frontier open-weight Large Language Models (LLMs) and Vision-Language Models (VLMs), such as GLM, Qwen, Kimi, and DeepSeek. This includes managing real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs, which has resulted in significant cost and latency reductions (e.g., producing a billion embeddings at approximately 20x lower cost and serving visual models at roughly 72% lower cost). The team also manages core platform components like the LLM Gateway, Agent Gateway, evaluation infrastructure, guardrails, and cost attribution. You will join a small, high-leverage team focused on building production infrastructure for Generative AI at DoorDash. Your primary focus will be on our open-weights model platform, covering inference and fine-tuning, including real-time GPU serving, high-throughput batch inference, and model fine-tuning. You will engage with various aspects of the platform such as model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is well-suited for an engineer who thrives on optimizing the cost and performance of GPU inference and fine-tuning in a rapidly evolving technical landscape where product requirements, model capabilities, vendor ecosystems, and cost/performance trade-offs are constantly changing.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior