DigitalOcean is seeking a Principal Engineer to own the model optimization discipline end-to-end for serving models across their entire accelerator fleet, including NVIDIA and AMD. This role is crucial for the Inference Platform team, which runs frontier open models in production on a heterogeneous GPU fleet. The goal is to close the gap between a model that merely runs and one that runs efficiently, saving millions of dollars in GPU time and ensuring customer Service Level Objectives (SLOs) are met. The engineer will be responsible for quantization strategy, kernel selection and authoring, attention and MoE execution, speculative decoding, and parallelism layout, adapting to changing model architectures and diverse vendor stacks (CUDA/Hopper/Blackwell and ROCm/MI300-class). The challenge lies in building a methodology, tooling, and upstream relationships to make every model fast on every GPU, quickly enough to support new model releases.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal
Education Level
No Education Listed