Member of Technical Staff, Sandbox Infrastructure

San Francisco Tensor CompanySan Francisco, CA
$275,000 - $315,000Onsite

About The Position

At SF Tensor, we're building the future of high-performance compute. We firmly believe that the future of AI depends on rethinking and rebuilding the stack from the hardware to the cloud. We aim to make compute faster, cheaper, and more available by developing our Kernel Optimizer, which finds the fastest code form for any vendor and cluster topology, and the Model Foundry, which manages runs, simplifies research, and moves workloads across clouds and chips based on price and availability. We are building the fastest GPU compiler in the world. Our approach allows us to search a wider space of code transformations while still guaranteeing correctness. To support this, we need to run an enormous amount of untrusted, freshly generated kernels on real silicon quickly and safely. This role is to build the serverless GPU container service that makes this possible across NVIDIA, AMD, TPU, and Trainium, at a scale and fidelity not available off-the-shelf. This service is critical for our compiler's measurement process, our internal post-training runs, and for providing isolated environments for customer workloads.

Requirements

  • Strong low-level systems engineering background: Linux kernel internals, containers, namespaces, cgroups, syscall interception or hypervisors
  • Experience in GPU systems engineering: drivers, runtimes or scheduling on accelerator fleets
  • Comfortable with distributed systems failure modes: preemption, partial failure, checkpoint/restore and dealing with states you can't afford to lose
  • Proficient in Go, C/C++ or Rust
  • Strong bias toward building the thing yourself when no vendor supports what you need

Nice To Haves

  • Worked directly with gVisor, Firecracker, Kata, QEMU/KVM or similar
  • Worked directly on CRIU, live migration or connection-preserving failover work
  • Familiar with NCCL/RCCL, RDMA, InfiniBand or vendor interconnects
  • Run large fleets on spot or other preemptible capacity
  • Familiar with bare-metal provisioning, hypervisors or fleet management at scale
  • Security background in isolation boundaries and untrusted code execution

Responsibilities

  • Extend our sandboxing stack to new vendors and accelerators
  • Build and maintain GPU virtualization below the runtime, including gVisor work at the driver and ioctl level
  • Make sandboxes first-class citizens on spot capacity, which means preemption-aware scheduling, checkpointing and rescheduling
  • Support multi-GPU and multi-node sandboxes, including the interconnect (NVLink, NVSwitch) and RDMA paths (InfiniBand, RoCE) paths those require
  • Own live migration end-to-end, including our socket-preserving migration
  • Guarantee measurement and profiling fidelity as well as their isolation
  • Work directly with the compiler, post-training and kernel teams to ensure their throughput is not capped by sandboxes

Benefits

  • Relocation assistance
  • Meaningful equity
  • Health insurance
  • Dental insurance
  • Vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service