Technical Director, Large-Scale AI Model Inferencing

Samsung SemiconductorSan Jose, CA
$219,000 - $351,000Onsite

About The Position

Inference is becoming a memory-bandwidth business. As models scale past what any single GPU can hold — KV caches grow with context, MoE expert weights spill beyond HBM, and new architectures change the rules of what "model state" even means — the winners will be the companies that treat memory as the core product of AI inference, not an afterthought. We are looking for a Hands-on Principal Engineer who combines deep, first-principles knowledge of AI model architectures (dense Transformers, Mixture-of-Experts, State Space Models, and hybrids) with production-scale inference expertise, to own the requirement for full-stack AI memory solutions at scale — spanning GPU HBM, host DRAM, CXL-attached memory pools, and NVMe/SSD tiers and Samsung Cognos, AI memory software that moves model state intelligently across them. This person will be the technical authority who connects model behavior to memory-system design: someone who can explain why an MoE router's activation pattern dictates an LRU expert cache policy, why a Mamba state cache breaks the assumptions of PagedAttention, and why disaggregated prefill/decode changes the required memory bandwidth per token by an order of magnitude — and then build the products that exploit those facts.

Requirements

  • BS in Computer/Electrical/Electronic Engineering or Computer Science, and 20 years of relevant experience MS in Computer/Electrical/Electronic Engineering or Computer Science with 18 years of relevant experience preferred.
  • 12+ years in systems engineering, with 4+ years hands-on in large-scale LLM inference or GPU systems performance — you have personally profiled, diagnosed, and fixed memory bottlenecks in production serving, not just read about them.
  • First-principles understanding of transformer-class model internals: you can derive KV-cache size formulas from attention math, explain MQA/GQA/MLA tradeoffs, and reason about activation-memory peaks during prefill.
  • Working expertise with MoE model behavior: routing, expert parallelism, load skew, and the weight-memory economics of serving models larger than GPU capacity.
  • Direct experience with at least one major inference stack's memory-management internals (vLLM PagedAttention/block manager, SGLang HiCache/token pools, TensorRT-LLM KV manager, or llama.cpp compute buffers) — code-level, not configuration-level.
  • Strong performance-engineering skills: bandwidth-bound vs. compute-bound analysis, NUMA and PCIe topology reasoning, RDMA basics, and fluency with GPU/CPU profilers.
  • Track record of building systems software at the memory/storage/IO layer — caches, tiering, paging, or storage engines — with production deployments.
  • Ability to write models and simulators, not just measure: analytical queueing, cache-hit-rate, and bandwidth models that predict system behavior before hardware exists.
  • Excellent written and verbal communication, including executive-level technical narrative; comfort being the technical face of the company in front of customers.

Nice To Haves

  • Experience with State Space Model or hybrid SSM-attention serving (Mamba-class state management, cache swapping for recurrent models) — rare and highly valued.
  • Contributions to open-source inference/caching projects (vLLM, SGLang, LMCache, HiCache, Mooncake, KTransformers, llama.cpp).
  • Experience with CXL memory pooling, CXL-attached tiering, or near-memory processing in real deployments or serious prototypes.
  • Experience with SSD/NVMe as a KV or expert cache tier, including QoS engineering for inference-grade latency.
  • Background in memory/storage product companies bringing hardware-software co-designed solutions to market.
  • You’re inclusive, adapting your style to the situation and diverse global norms of our people.
  • An avid learner, you approach challenges with curiosity and resilience, seeking data to help build understanding.
  • You’re collaborative, building relationships, humbly offering support and openly welcoming approaches.
  • Innovative and creative, you proactively explore new ideas and adapt quickly to change.

Responsibilities

  • Serve as expert on how different model families consume and move memory, and translate that into memory-product requirements.
  • Model the memory footprint, bandwidth demand, and access patterns of frontier open-weight models (e.g., Llama/Qwen-class dense, DeepSeek/Kimi-class MoE, Jamba-class hybrids) and publish internal reference architectures for each.
  • Track the model landscape as a roadmap input: anticipate what coming architectures (longer contexts, agentic multi-session reuse, reasoning-loop workloads, speculative decoding drafts) will demand from memory systems 12–24 months out.
  • Own deep expertise in production inference stacks — SGLang (HiCache), vLLM (PagedAttention, LMCache integration), NVIDIA Dynamo, TensorRT-LLM, llama.cpp-class engines — including their memory-management internals, not just their flags.
  • Drive inference performance engineering: continuous batching, chunked prefill, disaggregated prefill/decode, prefix and radix caching, speculative decoding, CUDA Graphs, and their interactions with memory tiering.
  • Own the latency/throughput/cost envelope: TTFT and TBT/TPOT SLOs, tokens-per-second per dollar, GPU memory utilization as the binding constraint, and the tradeoff curves between cache hit rate, memory capacity, and bandwidth.
  • Define benchmarking and characterization methodology: realistic agentic and long-context workloads (multi-turn reuse, session persistence, RAG prefixes), KV-cache reuse-rate measurement, and bandwidth-latency profiling across the full hierarchy (Nsight, PyTorch Profiler, vendor memory tools).
  • Define engineering requirements, with proof, for tiered memory systems for inference at fleet scale: HBM as L1, host DRAM (pinned, NUMA-aware pools) as L2, CXL-attached memory pools as an elastic tier, and NVMe/SSD as capacity tier — with the policies (admission, eviction, prefetch, placement) that make the hierarchy behave like one memory.
  • Design expert-weight offloading solutions for MoE serving: host-resident expert pools, GPU-resident expert caches with bandwidth-adaptive fill/evict policies, and CPU/CXL-execution hybrid paths — informed by the routing statistics of real models.
  • Translate model knowledge into product: write the requirements, reference architectures, and performance models that guide memory hardware and firmware roadmaps (HBM capacity/bandwidth, CXL device behavior, SSD QoS for cache tiers), and validate with end-to-end prototypes on real inference workloads.
  • Develop and Deliver POCs: demos and published benchmarks showing inference TCO improvement from the memory stack — e.g., context capacity multiplied at constant GPU count, or cost-per-token reduced through cache-hit-rate gains — credible to both CTOs and PhD researchers.
  • Set multi-year technical strategy for AI memory solutions; own build-vs-adopt-vs-contribute decisions across the open-source inference and caching ecosystem (vLLM, SGLang, LMCache, Cognos-style KV stores) and drive upstream contributions where strategic.
  • Lead architecture reviews and deep-dive design sessions; write the documents that become the company's standard for how we talk about memory for AI.
  • Represent the company with customers and partners at the deepest technical level: serve as the expert voice in CTO-to-CTO conversations, design wins, and standards discussions.
  • Mentor senior engineers and grow a bench of architecture talent across the model-to-memory boundary.

Benefits

  • Medical/Dental/Vision/401k
  • Charitable giving match
  • 4+ weeks of paid time off a year, plus holidays and sick leave
  • Stipend for fertility care or adoption
  • Medical travel support
  • Virtual vet care
  • On-demand apps and free confidential therapy sessions
  • Onsite Café and gym, plus virtual classes
  • Flexible environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service