Member of Technical Staff, GPU Kernels

San Francisco Tensor CompanySan Francisco, CA
$285,000 - $315,000Onsite

About The Position

At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the compiler to the cloud. We aim to make compute faster, cheaper, and more available by eliminating the friction between these components. Our Kernel Optimizer automatically finds the fastest form of code for any vendor and cluster topology, and our Model Foundry manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are backed by prominent investors and are looking for individuals who believe that a leap in compute is necessary for the next leap in AI. This role focuses on GPU Kernel Engineering, pushing the boundaries of what hardware can achieve. You will write and hand-optimize kernels for pre-training, post-training, and inference across various hardware platforms (NVIDIA, AMD, TPU, Trainium). Your learnings will be translated into structures for the compiler's search space, including instruction sequences, scheduling strategies, and cost signals. You will have access to advanced tooling due to our ownership of the stack down to the ISA, enabling the creation of kernels not expressible in standard CUDA. Our team possesses deep hardware understanding, including a bit-exact software model of Blackwell's tcgen05. You will contribute to significant projects, such as improving AlphaFold v3 throughput by 3.4x, post-training robotics models on Trainium, and running custom RL engines at scale.

Requirements

  • Track record of hand-writing kernels that match or beat vendor libraries.
  • Comfortable reading PTX, SASS, GCN/CDNA ISA, or equivalent machine-level assembly.
  • Fluent with low-level profiling tools like Nsight Compute, Nsight Systems, rocprof, omniperf, or their equivalents.
  • Solid systems programming skills in C++ and CUDA or ROCm/HIP.
  • Working understanding of how high-level ML operations map onto hardware, including framework layer inefficiencies.

Nice To Haves

  • Experience with compiler backends (MLIR, LLVM, codegen, instruction selection, scheduling).
  • Research experience in superoptimization, program synthesis, formal verification, or search-based compilation.
  • Experience with silicon beyond NVIDIA (AMD MI-series, TPU, Trainium) or mobile edge GPUs (Metal, Mali, Adreno).
  • Experience with distributed AI training at depth.
  • Experience in high-speed interconnects (NVLink, NVSwitch, InfiniBand, RoCE).
  • HPC background: large-scale scientific computing, MPI, supercomputing.
  • Experience in driver development or background in EE, computer architecture, or hardware design.

Responsibilities

  • Write and hand-optimize kernels for real workloads to find the performance ceiling before automated search.
  • Profile at the microarchitectural level, analyzing SM and CU utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput, and occupancy tradeoffs.
  • Debug issues down to clock behavior, thermal throttling, and driver paths.
  • Translate hard-won knowledge into machine-searchable structures for the compiler's exploration.
  • Work below PTX at the ISA-level, reasoning about SASS and cubins to emit schedules not expressible in PTX.
  • Build performance models, microbenchmarks, and tooling to predict kernel behavior.
  • Collaborate with the formal correctness team to ensure aggressive kernels ship with formal proofs.

Benefits

  • Meaningful equity
  • Relocation assistance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service