Software Engineer, ML Serving Platform

DoorDash USA•Sunnyvale, CA
•$130,600 - $192,000•Remote

About The Position

DoorDash’s ML Serving Platform delivers tens of millions of predictions per second, powering search, recommendations, advertising, delivery estimates, and logistics across DoorDash, Wolt, and Deliveroo. Our customers are modelers and engineering teams across our internal business verticals. We build self-serve infrastructure that empowers them to bring new models into production, adopt open source software and models, and expand what they can accomplish with machine learning at scale. We work at the intersection of a rapidly evolving open source ecosystem and demanding production workloads. With active customer demand and growing modeling ambitions, we’re advancing the infrastructure behind today’s predictions while building the capabilities that enable the next generation of ML innovation. You’ll help build the next generation of our ML serving platform, connecting request routing and online feature retrieval with model inference on CPU and GPU infrastructure. You’ll tackle challenging infrastructure problems involving latency, reliability, resource efficiency, and scale, and make those capabilities accessible through self-serve tools and workflows. Working alongside experienced platform engineers, you’ll own defined projects from technical design and implementation through testing, rollout, and production support. You’ll partner directly with internal teams to understand emerging modeling requirements, remove adoption barriers, and turn advances in open source technology into measurable production improvements.

Requirements

  • 2+ years of software engineering experience building and maintaining production services or infrastructure.
  • Strong computer science fundamentals and proficiency in a backend or systems programming language such as Java, Kotlin, Go, C++, or Python.
  • Understand distributed systems fundamentals, including concurrency, networking, timeouts, failure handling, and performance trade-offs.
  • Enjoy challenging infrastructure problems and can independently turn a defined problem into a technical design, tested implementation, and safe production rollout.
  • Experience debugging production systems and using operational data to improve reliability, performance, or cost.
  • Care about the engineers using your platform and collaborate effectively with customers and partners to make complex infrastructure easier to adopt.
  • Hold a degree in Computer Science or a related field, or have equivalent practical experience.

Nice To Haves

  • Experience with ML inference infrastructure, online feature retrieval, or other latency-sensitive distributed services.
  • Experience operating containerized workloads on Kubernetes, including deployment, resource management, or autoscaling.
  • Experience building self-serve developer platforms, APIs, or automation that helps other engineers move faster.
  • Experience integrating open source infrastructure or models and validating their behavior under production workloads.
  • Familiarity with CPU/GPU performance profiling, inference runtimes, or model deployment workflows.

Responsibilities

  • Empower modelers through self-serve infrastructure. Build tools, APIs, and workflows that help internal teams deploy, configure, validate, and operate models independently, accelerating ML adoption across business verticals.
  • Bring open source innovation into production. Evaluate and integrate evolving inference frameworks and enable open source models, translating promising capabilities into reliable, efficient services at scale.
  • Solve demanding inference infrastructure problems. Improve the systems that route prediction requests, retrieve online features, and execute models on a platform serving tens of millions of predictions per second.
  • Help evolve our disaggregated serving architecture. Build modular components that allow routing, feature retrieval, and model execution to evolve and scale independently as workloads and modeling requirements change.
  • Improve Kubernetes deployment and autoscaling. Help workloads respond to changing traffic while meeting latency and availability requirements and using CPU and GPU resources efficiently.
  • Make adoption and rollout easier. Build integrations, validation, and migration tooling that help teams adopt new serving capabilities and our unified global platform with confidence.
  • Own performance and reliability in production. Use benchmarking, profiling, metrics, and tracing to identify bottlenecks; participate in on-call and improve automation and runbooks to make the platform easier to operate.

Benefits

  • 401(k) plan with employer matching
  • 16 weeks of paid parental leave
  • Wellness benefits
  • Commuter benefits match
  • Paid time off
  • Paid sick leave
  • Medical, dental, and vision benefits
  • 11 paid holidays
  • Disability and basic life insurance
  • Family-forming assistance
  • Mental health program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service