Software Engineer, High Performance Computing

EventualSan Francisco, CA
Onsite

About The Position

Eventual is building a distributed data engine purpose-built for multimodal AI, called Daft. Our open-source engine is already processing petabytes of data daily at major companies like Amazon and FAANG, and is in production at leading AI companies. We are developing a video-native index on top of our engine to stream curated datasets to GPUs at line rate, aiming to saturate the latest hardware like B200s and future systems like NVL72 and Vera Rubin. We are partnering with top Physical AI labs and public AI infrastructure companies. Founded in 2022, Eventual has raised $30M from prominent investors and has a world-class team from companies like AWS, Render, Pinecone, and Tesla. We are looking for passionate individuals to join our small, powerful team working together 4 days a week in our SF Mission district office.

Requirements

  • Obsession with systems-level performance. You can recite Jeff Dean's "numbers every programmer should know" in your sleep. You eat flamegraphs for breakfast.
  • Strong opinions on io_uring — love it or hate it, you've earned the opinion.
  • Live and breathe Rust, C++, or C. You reach for them when it matters and you know why.
  • Strong familiarity with operating systems — page cache, scheduling, syscalls, NUMA, memory hierarchies.
  • A sense for where bytes actually go: NVMe vs. memory vs. network vs. PCIe vs. NVLink, and the throughput and latency budgets of each.

Nice To Haves

  • Experience working with GPUs is a plus, but you don't need it on day one.
  • Experience working with SLURM, Kubernetes for GPU workloads, or other HPC schedulers.
  • Hands-on CUDA experience.
  • Deep expertise on memory and caching subsystems — page cache tuning, hugepages, NUMA pinning, GPU-Direct Storage.
  • Worked on video decode pipelines (PyAV, decord, NVDEC) or PyTorch DataLoader internals.
  • Contributed to open-source systems projects in Rust/C++.

Responsibilities

  • Design and build the video-native dataloader: rank-aware, NVMe-cached, random-access into clips, returns tensors directly to the GPU.
  • Profile and optimize the full data path from object store → NVMe → page cache → host RAM → device RAM.
  • Eliminate every avoidable copy and stall.
  • Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs.
  • Push toward Vera Rubin bandwidth requirements.
  • Own performance benchmarks against customer baselines (custom DataLoaders, DALI, decord, LeRobot) and against our own historical numbers — regressions get caught at PR time.
  • Partner with researchers at our partner labs to land the loader in their training stack and measure MFU end-to-end.
  • Work cross-team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model-output ingestion path.

Benefits

  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team-building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment.
  • 401(k) plan with match.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service