Software Engineer, High Performance Computing

EventualSan Francisco, CA
$150,000 - $250,000Onsite

About The Position

Eventual is building a distributed data engine purpose-built for multimodal AI, called Daft. Their open-source engine is already running at scale for major companies like Amazon and other FAANG companies, and is in production at Mobileye, TogetherAI, and CloudKitchens. They are developing a video-native index on top of their engine for Physical AI, designed to stream curated datasets to GPUs at high speeds, aiming to saturate current and future high-end GPU hardware. The company was founded in 2022, has raised $30M from prominent investors, and has a team with experience from leading tech companies like AWS, Render, Pinecone, and Tesla. They are looking for individuals passionate about solving data loading challenges in AI training.

Requirements

  • Obsession with systems-level performance. You can recite Jeff Dean's "numbers every programmer should know" in your sleep. You eat flamegraphs for breakfast.
  • Strong opinions on io_uring — love it or hate it, you've earned the opinion.
  • Live and breathe Rust, C++, or C. You reach for them when it matters and you know why.
  • Strong familiarity with operating systems — page cache, scheduling, syscalls, NUMA, memory hierarchies.
  • A sense for where bytes actually go: NVMe vs. memory vs. network vs. PCIe vs. NVLink, and the throughput and latency budgets of each.

Nice To Haves

  • Experience working with GPUs is a plus, but you don't need it on day one.
  • Experience working with SLURM, Kubernetes for GPU workloads, or other HPC schedulers.
  • Hands-on CUDA experience.
  • Deep expertise on memory and caching subsystems — page cache tuning, hugepages, NUMA pinning, GPU-Direct Storage.
  • Worked on video decode pipelines (PyAV, decord, NVDEC) or PyTorch DataLoader internals.
  • Contributed to open-source systems projects in Rust/C++.

Responsibilities

  • Design and build the video-native dataloader: rank-aware, NVMe-cached, random-access into clips, returns tensors directly to the GPU.
  • Profile and optimize the full data path from object store → NVMe → page cache → host RAM → device RAM. Eliminate every avoidable copy and stall.
  • Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs. Push toward Vera Rubin bandwidth requirements.
  • Own performance benchmarks against customer baselines (custom DataLoaders, DALI, decord, LeRobot) and against our own historical numbers — regressions get caught at PR time.
  • Partner with researchers at our partner labs to land the loader in their training stack and measure MFU end-to-end.
  • Work cross-team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model-output ingestion path.

Benefits

  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team-building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment.
  • 401(k) plan with match.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service