Principal Engineer, Inference Memory and Storage Systems

DigitalOcean•Seattle, WA
•$249,600 - $312,000•Hybrid

About The Position

DigitalOcean is seeking a Principal Engineer to lead the technical vision and roadmap for the memory and storage layer of its inference platform. This role is crucial for optimizing the serving stack of frontier open models on GPU infrastructure at scale. The goal is to create a unified, multi-tier memory and storage substrate that spans GPU HBM, host DRAM, local NVMe, and remote object storage, accessible through a consistent interface to all serving engines and routers. The selected candidate will guide teams, contribute to open-source projects, and translate complex systems work into tangible customer benefits.

Requirements

  • 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure, with a proven track record of delivering and operating production services
  • Deep understanding of memory hierarchies (GPU HBM, host DRAM, NVMe, remote/object storage) and experience designing systems that span tiers for performance and cost
  • Experience with distributed caching or key-value systems, optimized for low latency under high concurrency
  • Hands-on experience with networked I/O and technologies like RDMA, NVMe-oF, or NVLink, and familiarity with aggregated and disaggregated deployment topologies for AI clusters
  • Strong systems programming skills in C/C++, Go, Rust, or Python, and ability to modify serving-engine internals
  • Proficiency in profiling and optimization across CPU, GPU, memory, and network, using measurement to drive architectural decisions and validate improvements
  • Excellent written and verbal communication skills
  • History of leading cross-functional efforts with product, infrastructure, and customer-facing teams

Nice To Haves

  • Contributions to open-source LLM serving or inference infrastructure projects (e.g., vLLM, SGLang, llm-d, NVIDIA Dynamo, LMCache), particularly related to KV cache offload, compression, or reuse
  • Experience designing a unified memory or storage layer presenting a single logical KV or object model across GPU, host, SSD, and cloud tiers in a hyperscale or public cloud environment
  • Familiarity with Kubernetes-based GPU orchestration, including DRA, MIG/MPS partitioning, and gateway/inference-extension routing patterns
  • Publications or patents in LLM systems, memory-disaggregated architectures, RDMA-based data planes, or CDN-like caching systems for ML workloads

Responsibilities

  • Defining and evolving a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, local NVMe tiers, and remote object storage for large-scale LLM inference
  • Architecting deep integrations with serving engines (vLLM, SGLang, TensorRT-LLM), focusing on KV cache offload, reuse, eviction policy, and cross-node sharing
  • Designing interfaces and protocols for disaggregated prefill/decode, peer-to-peer KV cache transfer, and cache-aware routing
  • Owning the eviction and admission strategy across memory tiers and defining metrics to validate policy effectiveness
  • Partnering with GPU infrastructure, networking, and platform teams to leverage technologies like GPUDirect, RDMA, NVMe-oF, and NVLink for low-latency cache access
  • Driving unit economics by modeling and validating the impact of cache strategies on cost and performance, and communicating these tradeoffs to product and pricing teams
  • Setting technical direction, improving standards through design reviews, mentoring engineers, and sponsoring roadmap initiatives
  • Representing DigitalOcean externally through open-source contributions, conference presentations, and customer-facing technical discussions

Benefits

  • Competitive array of benefits
  • Employee Assistance Program
  • Local Employee Meetups
  • Flexible time off policy
  • Reimbursement for relevant conferences, training, and education
  • Access to LinkedIn Learning's 10,000+ courses
  • Bonus opportunities based on company and individual performance
  • Equity compensation, including equity grants upon hire
  • Option to participate in Employee Stock Purchase Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service