Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is a critical part of the platform, encompassing fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes it. This role is within a platform engineering team, requiring both deep network expertise and software development capabilities. The ideal candidate will have experience with large-scale networks, as documentation alone is insufficient for understanding a 20,000 GPU training cluster. Production code development is essential, as manual CLI changes are not reliable at this scale. Configuration is model-driven, validated in CI, and applied by tooling. This role involves working alongside platform engineers in the same repositories and adhering to the same standards.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed