Principal Engineer, Model Optimizations

DigitalOcean•Seattle, WA
•Hybrid

About The Position

DigitalOcean is seeking a Principal Engineer to own the model optimization discipline end-to-end for serving models across their entire accelerator fleet, including NVIDIA and AMD. This role is crucial for the Inference Platform team, which runs frontier open models in production on a heterogeneous GPU fleet. The goal is to close the gap between a model that merely runs and one that runs efficiently, saving millions of dollars in GPU time and ensuring customer Service Level Objectives (SLOs) are met. The engineer will be responsible for quantization strategy, kernel selection and authoring, attention and MoE execution, speculative decoding, and parallelism layout, adapting to changing model architectures and diverse vendor stacks (CUDA/Hopper/Blackwell and ROCm/MI300-class). The challenge lies in building a methodology, tooling, and upstream relationships to make every model fast on every GPU, quickly enough to support new model releases.

Requirements

  • 12+ years in performance-critical systems, with substantial recent experience optimizing LLM inference in production.
  • Deep understanding of GPU architecture and the inference performance model, including memory bandwidth vs. compute bounds, arithmetic intensity, kernel launch and scheduling overhead, and differences between prefill and decode.
  • Hands-on kernel-level experience on at least one vendor stack (CUDA/CUTLASS/Triton or ROCm/HIP/CK) and the demonstrated ability and appetite to work across both.
  • Practical quantization expertise, including the judgment to assess when a technique that benchmarks well might fail a customer's accuracy bar.
  • Familiarity with the internals of at least one major serving engine (vLLM, SGLang, TensorRT-LLM), sufficient to contribute non-trivial changes upstream.
  • Strong Python and C++/CUDA skills.
  • Rigor in profiling using tools like Nsight, rocprof, and equivalents.
  • A measurement-first disposition, with claims backed by reproducible benchmarks and an understanding of how inference benchmarks can mislead.
  • Excellent written and verbal communication skills.
  • Experience leading cross-functional efforts spanning infrastructure, product, and customers.

Nice To Haves

  • Meaningful upstream contributions to vLLM, SGLang, TensorRT-LLM, PyTorch, Triton, or ROCm.
  • Experience bringing up a new accelerator family for production inference, including handling numerics differences, missing kernels, and immature libraries.
  • Experience with disaggregated prefill/decode serving and its interaction with parallelism and quantization choices.
  • Track record of same-week or day-zero enablement for newly released frontier open models.
  • Publications or patents in efficient inference, quantization, or GPU kernel design.

Responsibilities

  • Setting the technical strategy for model optimization across the fleet, including deciding which techniques to invest in, consume upstream, and how to make these decisions.
  • Owning quantization end-to-end, covering techniques like FP8, FP4/MXFP4, INT8, and weight-only schemes (AWQ, GPTQ), including calibration methodology, accuracy budgets, and evaluation gates.
  • Driving performance for modern architectures at the execution level, focusing on MoE routing, fused expert kernels, MLA and GQA attention variants, long-context and sliding-window attention, and their associated memory-movement patterns.
  • Leading speculative decoding efforts, including draft models, EAGLE/Medusa-class methods, n-gram and lookahead approaches, and tuning acceptance rates for practical benefit.
  • Writing and tuning kernels (CUDA, Triton, CUTLASS, and their ROCm counterparts like HIP, Composable Kernel, hipBLASLt, AITER) where upstream solutions are insufficient, and knowing when to avoid such efforts.
  • Establishing repeatable methods for choosing parallelism layouts (TP, PP, EP, attention DP) per model, GPU family, and traffic shape, moving away from tribal knowledge.
  • Building benchmarking and regression infrastructure to ensure the trustworthiness of optimization claims, including agent-shaped and long-context traffic distributions, TTFT/ITL/throughput measurement, and accuracy validation.
  • Making AMD a first-class target by driving upstream work in vLLM, SGLang, and TensorRT-LLM.
  • Partnering with NVIDIA and AMD engineering teams on pre-silicon enablement, early-access hardware, and roadmap feedback.
  • Setting technical direction, mentoring senior and staff engineers, and representing DigitalOcean in upstream communities, at conferences, and in customer technical deep dives.

Benefits

  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses.
  • Employee Assistance Program.
  • Local Employee Meetups.
  • Flexible time off policy.
  • Bonus in addition to base salary.
  • Equity compensation, including equity grants upon hire.
  • Option to participate in Employee Stock Purchase Program.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service