At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the compiler to the cloud. We aim to eliminate the friction that results in a tax on researchers, making compute faster, cheaper, and more available. Our Kernel Optimizer finds the fastest possible form of code for any vendor and cluster topology, and our Model Foundry manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are backed by prominent investors and are looking for individuals who believe that advancements in AI require advancements in compute first. We are hiring a Member of Technical Staff for GPU Compiler Engineering to build the machine that drives our search-based compiler. Our compiler holds the #1 spot on NVIDIA's kernel benchmark by proving correctness at the end, allowing for a wider search space. You will work on the entire pipeline, from StableHLO ingestion to backend code generation and executable binaries. You will have access to unique tooling due to our full-stack ownership, enabling us to create kernels others cannot express. Our custom LLVM backend emits cubins directly, allowing for deep optimization at the instruction level. The team's deep understanding of hardware is demonstrated by our bit-exact software model of Blackwell's tcgen05. You will contribute to significant performance improvements in areas like pre-training AlphaFold v3, post-training robotics models, and running large-scale RL engines.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed