Senior Network Engineer (3x Openings)

VoltaPalo Alto, CA
Onsite

About The Position

Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is a critical part of the platform, encompassing fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes it. This role is within a platform engineering team, requiring both deep network expertise and software development capabilities. The ideal candidate will have experience with large-scale networks, as documentation alone is insufficient for understanding a 20,000 GPU training cluster. Production code development is essential, as manual CLI changes are not reliable at this scale. Configuration is model-driven, validated in CI, and applied by tooling. This role involves working alongside platform engineers in the same repositories and adhering to the same standards.

Requirements

  • 4+ years in data center or cloud network engineering in production environments.
  • Deep understanding of Ethernet fabric, including leaf-spine design and BGP (including unnumbered BGP, ECMP).
  • Experience with EVPN and VXLAN in production, including multi-tenant load scenarios.
  • Production software development experience in Python or Go, with code managed in shared repositories under standard review, testing, and CI processes.
  • Experience with network automation, including configuration as code, source of truth systems (e.g., NetBox, Nautobot), and declarative/idempotent workflows.
  • Solid Linux fundamentals and command-line proficiency.
  • Familiarity with Kubernetes networking and its interaction with the underlying fabric.
  • Multi-vendor capability, able to work across major OEM platforms.
  • Willingness to be on-site during critical cluster bring-up periods.
  • Clear written communication skills for designs, decisions, and failure analysis.

Nice To Haves

  • Fluency with AI-assisted development tools (agentic CLI, IDE assistants, orchestrating coding agents).
  • Experience with RoCE v2 at scale, including PFC, ECN tuning, and DCQCN.
  • InfiniBand production experience (fat tree topology, UFM, fabric partitioning, adaptive routing, SHARP).
  • Experience with NVIDIA Spectrum-X, NetQ, Cumulus, SONiC, or whitebox platforms.
  • Experience with gNMI, OpenConfig, or NETCONF and YANG for configuration and telemetry.
  • Experience with IPv6 at production scale (dual stack design, v6 BGP peering, addressing architecture).
  • Experience with ASN operations (public ASNs, transit/IX peering, RPKI/IRR hygiene, DDoS posture).
  • Experience with OVN, OVS, SR-IOV, DPDK, or BlueField DPU based networking.
  • Network simulation or emulation experience (containerlab, NVIDIA Air).
  • Familiarity with NVLink and NVSwitch topologies and NCCL behavior.
  • Depth in Go or Rust beyond working proficiency.
  • Open source contributions to networking or infrastructure projects.
  • Experience working distributed across time zones with international counterparts.

Responsibilities

  • Design, build, and operate the network fabric across sites, including leaf-spine Ethernet, BGP underlay, EVPN and VXLAN overlay, and multi-tenant isolation.
  • Build and extend the software that manages the fabric, including configuration generation from a source of truth, validation pipelines, drift detection, and change management tooling.
  • Contribute production Python or Go code to the platform codebase, integrating the physical fabric with the IaaS control plane.
  • Own the exposure of network capabilities through APIs and abstractions for tenant networking.
  • Instrument the fabric with streaming telemetry, topology-aware metrics, and fabric understanding tools.
  • Collaborate with bring-up teams during cluster deployment, including fabric build, validation, acceptance testing, and translating pain points into platform features.
  • Debug complex network issues such as congestion, packet loss under collective communication load, and discrepancies between controller and hardware state.
  • Participate in on-call rotations, incident response, and follow-up work to address structural gaps.
  • Evaluate and challenge OEM and partner designs and configurations.
  • Participate in code reviews, technical design discussions, and cross-team collaboration in an Agile environment.
  • Specialize in areas such as GPU fabric (RoCE v2, InfiniBand), SDN and overlay integration, edge connectivity, or fabric observability, depending on background.

Benefits

  • Competitive salary
  • Equity in Volta
  • Retirement/pension contributions
  • Comprehensive health, wellbeing and insurance benefits
  • Generous number of vacation days
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service