At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the cloud. We aim to make compute faster, cheaper, and more available by developing our Kernel Optimizer, which finds the fastest code form for any vendor and cluster topology, and the Model Foundry, which manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are building the fastest GPU compiler in the world. Our approach allows us to search a wider space of code transformations while still guaranteeing correctness. To support this, we need to run an enormous amount of untrusted, freshly generated kernels on real silicon quickly and safely. This role is to build the serverless GPU container service that makes this possible across NVIDIA, AMD, TPU, and Trainium, at a scale and fidelity not available off-the-shelf. This service is critical for our compiler's measurement process, our internal post-training runs, and for providing isolated environments for customer workloads.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed