Member of Technical Staff

WaferSan Francisco, CA
Onsite

About The Position

Our mission at Wafer is to maximize intelligence per watt by using AI to optimize AI infrastructure, achieving orders of magnitude better energy and cost efficiency per token. We believe cheap intelligence is the most essential piece of technology for a future of abundance. We care about building a future where intelligence is "too cheap to meter." Wafer commercializes these efforts by serving serverless and dedicated inference for open source LLMs at the best performance per dollar. Our core bet is doing this through autonomous optimization of heterogeneous hardware.

Requirements

  • Infinitely Resourceful
  • Exceptionalism
  • Unreasonable Standards
  • Company Over Self
  • High EQ
  • Learns Quickly
  • First Principles Thinker
  • Complete autonomy of how to solve problems
  • Talking to customers
  • Writing custom GPU kernels in esoteric hardware

Responsibilities

  • Ship day-zero support for new open-source models, tuned for latency and throughput
  • Optimize the serving stack: batching, KV cache, speculative decoding, quantization
  • Write and tune kernels in CUDA, HIP, and Triton for NVIDIA, AMD, TPU, Trainium, D-Matrix, and more.
  • Design, deploy, and operate heterogeneous clusters across vendors
  • Run production inference across a mixed fleet: reliability, observability, and cost per token at scale

Benefits

  • $200K base salary + 1–2% equity
  • Fully covered medical, dental, and vision insurance
  • Daily lunch and dinner
  • Unlimited PTO
  • Parental leave
  • $1K/month housing stipend (post-tax) if you live within walking distance (0.5 miles) from the office
  • Covered Uber/Waymo from/to office
  • Visa sponsorship available
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service