Member of Technical Staff

Transparent Search Group•Seattle, WA
•Onsite

About The Position

Nuance Labs is hiring an experienced ML Infrastructure/Systems Engineer (2+ years) to own end-to-end inference optimization across LLMs, audio models and diffusion components, with a focus on latency, throughput and cost.

Requirements

  • 2+ years building and maintaining production ML systems
  • Designing scalable infrastructure from scratch
  • Track record optimizing latency, throughput and cost
  • Debugging distributed systems

Nice To Haves

  • Video or audio model experience
  • CUDA kernels and low-level optimization
  • Real-time video streaming (WebRTC)

Responsibilities

  • Own end-to-end inference optimization across the model stack.
  • Implement and tune KV cache strategies for long-context conversations.
  • Evaluate, deploy and extend serving frameworks such as vLLM, SGLang and TensorRT-LLM.
  • Profile and benchmark latency and throughput and remove bottlenecks.
  • Accelerate diffusion inference and apply quantization techniques (INT8, INT4, GPTQ, AWQ).
  • Build internal tooling that makes optimization work faster and more rigorous.

Benefits

  • HSA with about $2,000 annual company contribution
  • 15 days PTO plus public holidays
  • Company-wide office closure week
  • Relocation assistance available
  • Visa sponsorship available
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service