Principal Network Engineer

NscaleSeattle, WA

About The Position

Nscale is seeking a Principal Network Engineer to join their AI Infrastructure & High-Performance Networking team. This role involves designing, validating, and operating all networking services that support both internal management platforms and customer-facing cloud infrastructure. The responsibilities include managing high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and data center networking. The team also serves as a third/fourth-line escalation point for the support organization. As a Principal Network Engineer, you will be a senior technical authority, setting technical direction for low-latency, high-bandwidth InfiniBand and Ethernet networks, owning critical technical domains, and improving architecture, automation, operational rigor, and engineering standards. The role requires a blend of hands-on engineering and architectural influence, defining reference architectures, ensuring consistency across sites, leading technical decisions and escalations, and mentoring engineers while collaborating with various teams and vendors.

Requirements

  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data center environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM/UFM, and fabric orchestration.
  • Expert-level knowledge of modern data center routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures, with production experience on platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible, Git-based workflows, and modern infrastructure-as-code and pipeline tooling such as Terraform, GitLab CI, or GitHub Actions; you treat the network as code rather than managing devices by hand.
  • Deep design and engineering experience with firewall platforms such as Juniper SRX and/or Palo Alto, including security policy architecture, high-availability design, and multi-tenant segmentation.
  • Experience designing network telemetry and observability for high-throughput, performance-sensitive environments.
  • Proven ability to lead complex technical decisions and incidents across networking, systems, storage, and HPC/AI workload teams, with the judgment to balance performance, reliability, operability, and delivery velocity.
  • Demonstrated experience defining architecture, standards, and technical strategy beyond a single project or site, and influencing engineering teams without relying on formal authority.
  • Strong communication and mentoring skills, with the ability to make complex technical trade-offs clear to engineering leaders, operators, and cross-functional partners.
  • Hands-on, adaptable, and comfortable operating with high ownership in a fast-paced environment building next-generation infrastructure for ML scale-out.

Responsibilities

  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale, with tight integration to bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS, and establish reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure — firewalls, NAT, VPN, and security policy architecture — across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model, building and guiding Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence by leading complex escalations and root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs.
  • Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns.
  • Ensure the accuracy and reliability of source-of-truth network inventory and configuration data, with changes flowing through structured engineering and change-management practices.
  • Partner with deployment, data center operations, platform engineering, systems, storage, and vendors on new site delivery and platform evolution.
  • Act as a technical mentor and force multiplier across the team through architecture reviews, design reviews, incident leadership, documentation, and knowledge sharing.
  • Identify systemic risks and architectural gaps across sites and drive durable solutions that improve scalability, reliability, and operational simplicity.

Benefits

  • medical
  • dental
  • vision
  • flexible paid time off
  • parental leave
  • retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service