Senior Inference Optimization Engineer - Dragonfly Portfolio

DragonflyUnited States (Remote),
Remote

About The Position

This is an application to join the talent network of Dragonfly, a crypto-native Venture Capital and research firm. They are sourcing for a Senior Inference Optimization Engineer for one of their portfolio companies. This company is building privacy-first consumer AI infrastructure and is looking for someone to work on the bleeding edge of LLM inference performance, focusing on throughput, latency, and cost per token at scale. The role involves optimizing GPU infrastructure, benchmarking inference engines, optimizing load-balancing algorithms, evaluating new inference optimization techniques, and assessing emerging inference hardware.

Requirements

  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
  • Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go. C++/CUDA a strong plus

Nice To Haves

  • custom Triton kernels
  • diffusion/image model inference optimization
  • open-source inference framework contributions

Responsibilities

  • Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
  • Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Optimize multivariate inference load-balancing algorithms within the inference routing system
  • Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service