Inference Performance Engineer

adaptionSan Francisco, CA
Hybrid

About The Position

The role focuses on owning the cost and performance of the inference stack, ensuring efficient model serving as workloads, traffic, and hardware evolve. The engineer will collaborate with fleet operators while managing core performance aspects like caching, batching, quantization, decoding, and kernel-level optimization. The goal is to enhance throughput and latency without sacrificing reliability or model quality.

Requirements

  • 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Nice To Haves

  • Great teammates who make work feel lighter and aren't afraid to go out on a limb with bold ideas.
  • Adaptability.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Benefits

  • Flexible work: In-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
  • Adaption Passport: Annual travel stipend to explore a country you've never visited.
  • Lunch Stipend: Weekly meal allowance for take-out or grocery delivery.
  • Well-Being: Comprehensive medical benefits and generous paid time off.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service