Member of Technical Staff, GPU Compiler

San Francisco Tensor CompanySan Francisco, CA
$285,000 - $315,000Onsite

About The Position

At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the compiler to the cloud. We aim to eliminate the friction that results in a tax on researchers, making compute faster, cheaper, and more available. Our Kernel Optimizer finds the fastest possible form of code for any vendor and cluster topology, and our Model Foundry manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are backed by prominent investors and are looking for individuals who believe that advancements in AI require advancements in compute first. We are hiring a Member of Technical Staff for GPU Compiler Engineering to build the machine that drives our search-based compiler. Our compiler holds the #1 spot on NVIDIA's kernel benchmark by proving correctness at the end, allowing for a wider search space. You will work on the entire pipeline, from StableHLO ingestion to backend code generation and executable binaries. You will have access to unique tooling due to our full-stack ownership, enabling us to create kernels others cannot express. Our custom LLVM backend emits cubins directly, allowing for deep optimization at the instruction level. The team's deep understanding of hardware is demonstrated by our bit-exact software model of Blackwell's tcgen05. You will contribute to significant performance improvements in areas like pre-training AlphaFold v3, post-training robotics models, and running large-scale RL engines.

Requirements

  • Deep experience in compiler infrastructure (LLVM, MLIR or similar)
  • Strong background in GPU architecture and low-level optimization (CUDA, ROCm or similar)
  • Hands-on experience with at least one of: PTX/SASS, GCN/RDNA assembly or other GPU ISAs
  • Familiarity with ML compiler stacks (XLA, TVM, Triton, torch.compiler or similar)
  • Solid systems programming skills in C++ and/or Rust
  • Proven track record of building production-grade compiler infrastructure

Nice To Haves

  • Experience in autotuning or search-based optimization
  • Background in formal verification, proof assistants, or SMT solvers
  • Experience writing or maintaining an LLVM backend
  • Background in distributed systems or multi-device compilation
  • Contributions to open-source compiler projects
  • Familiarity with large-scale training infrastructure
  • Experience in (Stable)HLO

Responsibilities

  • Build and extend MLIR dialects and passes to optimize training and inference workloads
  • Work in our LLVM backend, below ptxas, on instruction selection, scheduling, register allocation, and direct cubin emission, and establish equivalent depth on AMD, TPU, and Trainium targets
  • Expand our search-based compiler infrastructure, including agent- and RL-driven program search and formal correctness proofs
  • Implement classic compiler optimizations tuned for large-scale training
  • Create hybrid codegen paths for cases where direct MLIR lowering isn't practical
  • Own testing, benchmarking, and performance regression systems, including bit-identical hardware models
  • Work closely with our research team and customer workloads to identify optimization opportunities

Benefits

  • Meaningful equity
  • Relocation assistance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service