Member of Technical Staff - Storage Infrastructure

Prime IntellectSan Francisco, CA
$150,000 - $300,000

About The Position

Prime Intellect is building the open superintelligence stack, providing infrastructure for AI labs. Their platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment for post-training at frontier scale. They are developing open-source models for long-horizon tasks and the platform to build them. Prime Intellect has raised $150M from prominent investors and individuals in the AI and infrastructure space. They are seeking individuals passionate about building at the intersection of frontier research, real infrastructure, and go-to-market for a nascent category.

Requirements

  • 3+ years building or operating production distributed storage systems.
  • Hands-on experience with at least one parallel or distributed filesystem or object storage platform (e.g., Lustre, BeeGFS, Ceph, or GPFS).
  • Strong Linux administration and performance troubleshooting skills.
  • Experience automating infrastructure operations in Python, Go, Bash, or similar languages.
  • Understanding of storage failure modes, data integrity, consistency, replication, and recovery.
  • Knowledge of block, file, and object storage semantics and their performance tradeoffs.
  • Familiarity with NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking.
  • Understanding of high-throughput storage networking and distributed client behavior.
  • Experience with capacity forecasting, observability, alerting, and safe maintenance procedures.
  • Knowledge of authentication, authorization, encryption, and secure data lifecycle management.

Nice To Haves

  • Experience supporting large GPU training clusters and high-volume checkpoint workloads.
  • Experience with S3-compatible object storage, data tiering, or distributed caching.
  • RDMA-enabled storage or GPUDirect Storage experience.
  • Kubernetes storage integrations or SLURM environments.
  • Experience with storage cost optimization.
  • Contributions to open-source storage systems.

Responsibilities

  • Build and operate storage systems for frontier AI workloads, ensuring reliable, high-throughput access to datasets, checkpoints, and model artifacts.
  • Balance performance, durability, availability, and cost as GPU clusters scale.
  • Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows.
  • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads.
  • Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads.
  • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services.
  • Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets.
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices.
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks.
  • Collaborate with compute and networking teams.

Benefits

  • Cash compensation range of $150,000–$300,000
  • Equity incentives
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service