Senior Director, Inference Products and Optimizations

DigitalOceanSan Francisco, CA
Remote

About The Position

DigitalOcean is seeking an experienced Senior Director of Engineering to lead a high-performing engineering team building and scaling its Large Language Model (LLM) inference products. This role involves overseeing the control plane, model optimization, and model architecture layers, aiming to bring simplicity to optimized LLM inference. The position owns DigitalOcean's inference product suite, including Serverless Inference, Dedicated Inference, Inference Router, Batch Inference, and Multimodal Inference, as well as the underlying model optimization and architecture stack. The goal is to deliver robust, cost-efficient systems that serve millions of users globally at scale and high performance.

Requirements

  • 10+ years of software engineering experience.
  • 6+ years in a technical leadership or management role, ideally within Inference Systems or AI/ML systems.
  • Deep expertise in distributed systems design, modern AI/ML technologies, Kubernetes at scale, and LLM inference.
  • Expertise in AI workload orchestration, scheduling, and resource management.
  • Ability to engage in deep technical discussions regarding scalable control plane design, inference engines (vLLM, SGLang), and model architectures.
  • Strong understanding of cloud-native multi-region architectures, microservices, and distributed systems fundamentals.
  • Strategic knowledge of GPU architectures (NVIDIA and/or AMD), interconnects (like NVLink), and hardware topology and their impact on AI performance.
  • Familiarity with concepts in container runtime internals, system isolation, and security contexts.
  • Expertise in defining, tracking, and operationalizing deep infrastructure and inference metrics (e.g., TTFT, TPOT) to drive performance improvements and meet service level objectives.
  • Demonstrated ability to translate complex technical requirements into user-focused product features.
  • Excellent communication skills, with the ability to explain technical decisions to non-technical stakeholders and align diverse teams.
  • A strong sense of ownership and a proactive drive to identify and resolve issues.

Nice To Haves

  • Understanding of the balance between innovation and reliability.

Responsibilities

  • Recruit, mentor, and coach engineers, fostering a culture of ownership, technical excellence, and continuous improvement.
  • Work with Product teams to define and execute the Product roadmap for all of DigitalOcean’s Inference Products.
  • Lead the design and evolution of the inference serving stack, driving technical strategy across vLLM, SGLang, and LLM-D to optimize throughput, latency, and GPU utilization.
  • Architect the model-serving and optimization layer, including quantization, KV-cache management, speculative decoding, and disaggregated serving.
  • Collaborate with Product Management, other engineering teams, and stakeholders to align priorities, manage dependencies, and communicate progress.
  • Ensure the production health, stability, and on-call rotation of all systems to maintain customer SLAs.
  • Institutionalize benchmarking frameworks, observability, and auto-tuning capabilities.
  • Encourage contributions to open-source inference engines.

Benefits

  • Competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses.
  • Bonus in addition to base salary.
  • Equity compensation, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service