Member of Technical Staff

Transparent Search Group•San Francisco, CA
•Onsite

About The Position

Wafer is an AI cloud on a mission to maximize intelligence per watt, using AI to optimize AI infrastructure. It serves serverless and dedicated inference for open-source LLMs at the best performance per dollar, using autonomous agents to write and tune GPU kernels across heterogeneous hardware, and is deploying its own hardware in co-located data centers. Founded in 2025, Wafer went from $0 to $4M in revenue in 8 weeks. It has about 7 people and is backed by Fifty Years, Liquid 2 Ventures and prominent AI angel investors. Wafer is hiring a Member of Technical Staff (1-6 years) to work close to the hardware on its inference stack. On a small team with massive surface area, you will do everything from talking to customers (10-20% of the role) to writing custom GPU kernels for esoteric hardware, with full autonomy over how you solve problems and a hand in setting the direction of its inference serving infrastructure.

Requirements

  • 1-6 years of software engineering focused on systems close to the hardware
  • Deep understanding of computer architecture, memory hierarchy and OS internals
  • Backend, infrastructure or systems work at a top-tier tech company, hardware-adjacent company, quant firm or strong AI startup
  • CS degree from a top undergraduate program
  • Evidence of exceptionalism (competitions, rankings, standout projects)
  • Enthusiasm for AI agents and coding tools in your own workflow
  • High EQ and clear communication with customers
  • Ability to work on-site in San Francisco 5 days a week

Nice To Haves

  • New grads with strong internships are considered
  • Deploying or operating infrastructure at scale (data centers, clusters, GPU fleets)
  • ML inference optimization (quantization, batching, KV cache, speculative decoding)
  • CUDA, Triton or GPU programming; OS, architecture or compilers coursework

Responsibilities

  • Ship day-zero support for new open-source models, tuned for latency and throughput.
  • Optimize the serving stack: batching, KV cache, speculative decoding and quantization.
  • Write and tune kernels in CUDA, HIP and Triton for NVIDIA, AMD, TPU, Trainium and other accelerators.
  • Design, deploy and operate heterogeneous clusters across vendors.
  • Run production inference across a mixed fleet with strong reliability, observability and cost per token.

Benefits

  • $1,000/month housing stipend for anyone living within half a mile of the office.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service