Cloud Storage Integration Engineer (Storage / Image / Registry)

Bitdeer Technologies Group•San Jose, CA

About The Position

GPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud — covering performance, multi-tenancy, and reliability — while also owning the image / driver / registry pipeline on the node-delivery critical path.

Requirements

  • 3+ years (Senior 6+) in storage engineering or platform infrastructure, with hands-on distributed / parallel file system experience.
  • Strong understanding of distributed file system internals — data / metadata separation, replication, consistency models, POSIX vs object semantics.
  • Proven experience integrating and operating distributed storage in production (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python / Go, CI / CD).

Nice To Haves

  • HPC / AI storage or multi-region storage experience a strong plus.

Responsibilities

  • Design and integrate distributed / parallel file systems (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training / inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
  • Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
  • Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
  • Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service