GPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud — covering performance, multi-tenancy, and reliability — while also owning the image / driver / registry pipeline on the node-delivery critical path.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed