We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure. You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving. Build, operate, and scale next-generation GPU infrastructure: You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world’s most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads. Build the playbook, then grow into the platform: Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed