Principal HPC Network Engineer (remote in the US)

MirantisRemote, USA
Remote

About The Position

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Requirements

  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.
  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
  • Experience with network performance analysis and troubleshooting tools.
  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
  • Strong analytical and problem-solving skills.

Nice To Haves

  • Experience with large-scale HPC clusters or AI/ML infrastructure.
  • Knowledge of RDMA, MPI, and low-latency networking concepts.
  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

Responsibilities

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.
  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.
  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
  • Develop and maintain documentation for network architecture, configurations, and operational procedures.
  • Participate in on-call rotations and provide escalation support for critical incidents.
  • Lead or contribute to network upgrades, migrations, and new deployments.

Benefits

  • Professional development and training
  • Attend conferences and working groups
  • Company outings, happy hours, hackathons, and tech talks
  • Receive a competitive compensation package with a strong benefits plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service