About The Position

NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. We’re looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.
  • Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale.
  • Contributions to open-source projects involving a distributed or parallel file system.
  • You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.
  • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
  • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).
  • Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
  • Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers.
  • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
  • 100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.

Nice To Haves

  • Maintainers or sustained contributions to widely used public projects.
  • Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
  • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
  • Kubernetes and CSI driver development for storage.
  • Hands-on experience with SPDK, libfabric, or FUSE performance optimization.

Responsibilities

  • Contribute to open-source file systems.
  • Contribute code to open-source parallel and distributed file systems, and distributed object storage.
  • Upstream fixes and features, and engage directly with the upstream communities and maintainers.
  • Serve as a hands-on storage software lead.
  • Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it.
  • Make the final technical calls on storage deliveries against measurable targets.
  • Triage and troubleshoot at scale.
  • Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery.
  • Validate architecture and capabilities.
  • Validate storage architecture, capabilities, performance, and durability.
  • Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.
  • Recommend configuration, tuning, and guidelines.
  • Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them.
  • Partner broadly.
  • Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.
  • Work AI-first.
  • Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service