Senior Director, Inference Products and Optimizations

DigitalOceanSeattle, WA
Hybrid

About The Position

DigitalOcean is seeking an experienced Senior Director of Engineering to lead a high-performing engineering team building and scaling its Large Language Model (LLM) inference products. This role involves overseeing the control plane, model optimization, and model architecture layers, aiming to bring simplicity to optimized LLM inference. The position owns DigitalOcean's inference product suite, including Serverless Inference, Dedicated Inference, Inference Router, Batch Inference, and Multimodal Inference, as well as the underlying model optimization and architecture stack. The goal is to deliver robust, cost-efficient systems that serve millions of users globally at scale and high performance.

Requirements

  • 10+ years of software engineering experience.
  • 6+ years in a technical leadership or management role, ideally within Inference Systems or AI/ML systems.
  • Deep expertise in distributed systems design, modern AI/ML technologies, Kubernetes at scale, and LLM inference.
  • Expertise in AI workload orchestration, scheduling, and resource management.
  • Ability to engage in deep technical discussions regarding scalable control plane design, inference engines (vLLM, SGLang), and model architectures.
  • Strong understanding of cloud-native multi-region architectures, microservices, and distributed systems fundamentals.
  • Strategic knowledge of GPU architectures (NVIDIA and/or AMD), interconnects (like NVLink), and hardware topology and their impact on AI training and inference performance.
  • Familiarity with concepts in container runtime internals, system isolation, and security contexts.
  • Expertise in defining, tracking, and operationalizing deep infrastructure and inference metrics (e.g., TTFT, TPOT) to drive performance improvements and meet service level objectives.
  • Demonstrated ability to translate complex technical requirements into user-focused product features.
  • Understanding of the balance between innovation and reliability.
  • Excellent communication skills, with the ability to explain technical decisions to non-technical stakeholders and align diverse teams around a shared vision.
  • A strong sense of ownership and a proactive drive to identify and resolve issues preventing the team from delivering value.

Responsibilities

  • Recruit, mentor, and coach engineers, fostering a culture of ownership, technical excellence, and continuous improvement.
  • Work with Product teams to define and execute the Product roadmap for all of DigitalOcean’s Inference Products.
  • Lead the design and evolution of the inference serving stack, driving technical strategy across vLLM, SGLang, and LLM-D to optimize throughput, latency, and GPU utilization.
  • Architect the model-serving and optimization layer, spanning quantization, KV-cache management, speculative decoding, and disaggregated serving.
  • Collaborate with Product Management, other engineering teams, and key stakeholders to align priorities, manage dependencies, and communicate progress and risks.
  • Ensure the production health, stability, and on-call rotation of all systems to maintain customer SLAs.
  • Institutionalize benchmarking frameworks, observability, and auto-tuning capabilities.
  • Encourage contributions to open-source inference engines.

Benefits

  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses.
  • Employee Assistance Program.
  • Local Employee Meetups.
  • Flexible time off policy.
  • Bonus in addition to base salary.
  • Equity compensation, including equity grants upon hire.
  • Option to participate in our Employee Stock Purchase Program.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service