At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the compiler to the cloud. We aim to make compute faster, cheaper, and more available by eliminating the friction between these components. Our Kernel Optimizer automatically finds the fastest form of code for any vendor and cluster topology, and our Model Foundry manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are backed by prominent investors and are looking for individuals who believe that a leap in compute is necessary for the next leap in AI. This role focuses on GPU Kernel Engineering, pushing the boundaries of what hardware can achieve. You will write and hand-optimize kernels for pre-training, post-training, and inference across various hardware platforms (NVIDIA, AMD, TPU, Trainium). Your learnings will be translated into structures for the compiler's search space, including instruction sequences, scheduling strategies, and cost signals. You will have access to advanced tooling due to our ownership of the stack down to the ISA, enabling the creation of kernels not expressible in standard CUDA. Our team possesses deep hardware understanding, including a bit-exact software model of Blackwell's tcgen05. You will contribute to significant projects, such as improving AlphaFold v3 throughput by 3.4x, post-training robotics models on Trainium, and running custom RL engines at scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed