Data Center Network Architect

SPANSan Francisco, CA
$175,000 - $230,000Onsite

About The Position

SPAN is seeking a Network Architect to lead the design, deployment, and operation of the network for its distributed data center fleet (XFRA). This role involves architecting high-performance networks within GPU clusters, managing connectivity between compute sites and customers, and designing the wide-area network that links the fleet to cloud platforms and partners. The architect will translate workload requirements into practical designs, balancing performance, availability, security, deployment speed, and cost. This is a hands-on technical leadership position requiring equipment selection and configuration, ISP service qualification, deployment automation, production issue troubleshooting, and establishing scaling standards. Collaboration with compute, orchestration, security, deployment, and commercial teams is essential.

Requirements

  • 8+ years of relevant network engineering experience, or equivalent demonstrated expertise, including ownership of production architecture and hands-on implementation.
  • Strong experience with both data center networks and WAN or ISP connectivity across multiple sites.
  • Deep understanding of TCP/IP, BGP, IPv4/IPv6, routing, switching, segmentation, VPNs, firewalls, QoS, and high-availability design.
  • Practical understanding of fiber access and carrier service models, including how to validate real performance, availability, and commercial service commitments.
  • Experience designing and troubleshooting high-speed server networks, including NICs, optics, cabling, MTU, congestion, and oversubscription.
  • Ability to translate application traffic patterns and service objectives into network requirements, and explain performance, resilience, and cost tradeoffs clearly.
  • Hands-on Linux troubleshooting and network automation using Python, Ansible, or comparable tools, with version-controlled configurations and disciplined change management.
  • A record of taking ambiguous problems from design through deployment and production support, working effectively with engineering teams, vendors, and field operators.

Nice To Haves

  • Production experience with GPU clusters, AI inference, HPC, InfiniBand, RoCE, RDMA, or NCCL performance analysis.
  • Experience with EVPN/VXLAN, leaf-spine fabrics, cloud interconnects, internet peering, or operating an autonomous system.
  • Experience operating a large fleet of unattended edge sites, including zero-touch provisioning, carrier-grade NAT constraints, and remote recovery.
  • Familiarity with Kubernetes networking, multi-tenant compute platforms, and network-aware workload scheduling.
  • Experience qualifying carrier services and negotiating technical requirements across multiple providers or regions.

Responsibilities

  • Own the network architecture, developing reference designs for residential nodes and commercial GPU clusters, including routing, switching, security gateways, addressing, management access, physical connectivity, and capacity planning.
  • Engineer data center connectivity, designing north-south paths for customer traffic, modeling distribution, storage access, and fleet management, as well as east-west fabrics for server communication. Evaluate high-speed Ethernet, InfiniBand, and RoCE based on workload requirements.
  • Select and qualify fiber services, evaluating options like PON broadband, dedicated internet access, Ethernet private lines, wavelength services, and dark fiber. Assess performance metrics, service guarantees, and costs.
  • Build provider relationships with ISPs and carriers for service availability, technical requirements, provisioning, testing, and incident escalation. Verify physical path diversity and dependencies for redundant connectivity.
  • Optimize the distributed fleet by defining connection strategies (direct internet, regional aggregation, private interconnects) and evaluating transit, peering, cloud connectivity, traffic engineering, and failover options.
  • Partner with orchestration engineering to provide network signals for workload placement, traffic steering, and recovery, considering model transfers and distributed inference communication.
  • Build security and resilience into deployments, implementing tenant isolation, traffic separation, encrypted connectivity, access controls, and DDoS mitigation. Design remote recovery and out-of-band access for unattended sites.
  • Automate and operate at scale by building repeatable provisioning, configuration validation, staged changes, rollback, and inventory management. Establish telemetry, service objectives, alerting, runbooks, and acceptance tests, and lead network incident diagnosis.

Benefits

  • Competitive compensation + equity grants
  • Comprehensive benefits: 100% employee premiums for base plans on medical, dental, vision with options for additional coverage.
  • Parental leave up to twenty four (24) weeks depending on eligibility
  • Comfortable, sunny office space located near BART and Caltrain public transit
  • Strong focus on team building and company culture: Employee Resource Groups, monthly social events, SPANcakes recognition breakfast, lunch, and learns
  • Flexible hours, one holiday per month, and flexible time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service