Staff Machine Learning Engineer, AI Platform

AdobeSan Jose, CA
$172,500 - $306,625

About The Position

Adobe's AI and generative AI products, from Firefly to the intelligence built into Creative Cloud and Experience Cloud, run on a shared compute and inference platform. It's the infrastructure every internal ML team uses to train models and serve them to production at global scale. As a Staff Machine Learning Platform Engineer, you'll own core parts of that platform. That means extracting maximum utilization from large GPU fleets, moving models from experiment to production without a rewrite, and serving inference at low latency and high throughput under real traffic. You set technical direction rather than take tickets, and the architecture you define shapes how hundreds of engineers across Adobe train and ship AI.

Requirements

  • 7+ years building and operating large-scale platform, infrastructure, or distributed systems in production, with direct ownership of performance, scalability, and reliability.
  • Deep expertise in distributed systems and cloud infrastructure, including Kubernetes, containerized workloads, and operating large multi-node and multi-region clusters.
  • Strong programming ability in Python and at least one systems language (Go, C++, Rust, or Java).
  • A track record of designing systems that other engineers build on, making deliberate architectural tradeoffs and taking them from design to production at scale.
  • A bias for measurable outcomes (latency, throughput, utilization, reliability) and the collaboration skills to drive them across teams and stakeholders.

Nice To Haves

  • Experience with GPU or accelerator scheduling, performance tuning, or fleet management.
  • Familiarity with ML framework internals or distributed training (PyTorch, FSDP, DeepSpeed) or modern inference stacks (vLLM, TensorRT-LLM, Triton, Ray Serve).
  • Experience operating ML or data infrastructure at the scale of a major ML-driven product organization

Responsibilities

  • Own the architecture and roadmap for major components of the ML compute and inference platform, such as training orchestration, GPU scheduling and utilization, model serving, or the developer-facing surfaces ML teams build on.
  • Design and operate distributed systems that run large-scale training and low-latency, high-throughput inference reliably across thousands of accelerators.
  • Drive multi-tenancy, elasticity, and cost/utilization efficiency across a shared GPU fleet serving many teams with competing demands.
  • Build the paths that move a model from experiment to production without re-implementation, from packaging and registry through deployment and safe rollout.
  • Set engineering standards for reliability, observability, and performance, and raise the bar for how the platform is built and operated.
  • Partner with ML researchers and product teams to turn emerging workloads into first-class platform capabilities, and inform capacity and hardware strategy.
  • Provide technical leadership and mentorship across the platform organization.

Benefits

  • comprehensive benefits programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service