Senior Applied Scientist / Engineer, Training & Inference

AdobeSan Jose, CA
$164,000 - $313,300

About The Position

Adobe Applied Science & Machine Learning (ASML) is seeking a Senior Applied Scientist / Engineer, Training & Inference to play a critical role in closing the gap between research and production for Adobe's next-generation video and image foundation models. In this role, you will serve as a technical owner for the training-to-deployment pipeline for our video and multimodal generation models. Rather than focusing solely on model research or systems infrastructure in isolation, you will bridge both — bringing the hands-on training expertise and the inference and deployment depth needed to take large generative models from the research cluster to reliable, performant, and cost-efficient production. This role is ideal for those who excels at the full arc of model development — distributed training at scale, inference optimization, and the practical engineering required to deploy and operate models reliably in production.

Requirements

  • Master's or PhD in Computer Science, Electrical Engineering, AI/ML, or a related field, or equivalent practical experience.
  • Hands-on experience with large-scale distributed training using PyTorch (FSDP, Tensor Parallelism, Pipeline Parallelism) across multi-node GPU environments.
  • Proven experience optimizing and deploying large generative models for production — including serving infrastructure, latency/throughput tuning, and cost-aware deployment.
  • Proficiency in Python and PyTorch, with experience working in large shared codebases and contributing to production-critical ML systems.
  • Demonstrated ability to take models from training through deployment, navigating the practical engineering challenges of reliability, reproducibility, and operational scale.
  • Demonstrated ability to independently own end-to-end technical areas, drive cross-team execution, and deliver high-quality systems on which product teams depend.

Nice To Haves

  • Experience training and deploying video, image, or multimodal generative models (e.g., diffusion models, flow matching, video generation).
  • Familiarity with inference serving frameworks such as TensorRT, vLLM, or equivalent.
  • Experience with performance profiling and optimization for both training and inference workloads.
  • Track record of shipping generative AI models to production at scale.
  • Prior work in an applied research environment bridging ML and systems engineering.

Responsibilities

  • Own key components of the training-to-deployment pipeline — from distributed training execution through inference optimization, serving, and production handoff — ensuring models are delivered reliably, performantly, and cost-efficiently.
  • Implement and operate distributed training strategies including PyTorch FSDP, Tensor Parallelism, and Pipeline Parallelism across multi-node GPU environments, ensuring correctness, stability, and scalability for large video and multimodal models.
  • Design and optimize inference and serving systems for large generative models, with a focus on latency, throughput, and cost across deployment targets.
  • Reduce the gap between trained model checkpoints and reliable production deployments — owning the practical work of hardening, validating, and operationalizing models at scale.
  • Identify and address inefficiencies across the training and inference stack — memory, communication, scheduling, and execution orchestration — with a clear focus on GPU efficiency and cost targets.
  • Partner closely with applied researchers, ML engineers, and infrastructure teams to align training and inference systems with model architecture needs and product delivery timelines.

Benefits

  • comprehensive benefits programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service