Network Solutions Architect, AI Factory Services

Lenovo•Morrisville, NC
•Hybrid

About The Position

Lenovo seeks a highly experienced Network Solutions Architect to join the Hybrid Cloud Solutions and AI Offering Engineering team within SSG. This senior-level role designs, deploys, validates, and optimizes high-performance network infrastructure for Lenovo's AI Factory and GigaFactory service offerings, supporting both enterprise-scale and hyperscale GPU environments powered by NVIDIA Spectrum-X Ethernet and InfiniBand fabrics. The architect will develop network reference architectures, deployment runbooks, performance validation procedures, and field-ready engineering documentation consumed by Lenovo Professional Services and Managed Services teams globally. This includes GPU cluster fabric design, BlueField DPU architectures, multi-tenant network isolation, high-performance RDMA and RoCEv2 deployments, and production-scale AI infrastructure supporting large distributed training and inference workloads. The ideal candidate brings deep expertise in NVIDIA Spectrum-X, Quantum InfiniBand, and BlueField DPU technologies, along with hands-on experience architecting, deploying, validating, and troubleshooting large-scale GPU clusters. Experience supporting hyperscale AI infrastructure, high-density liquid-cooled environments, GPU cluster bring-up, infrastructure validation, performance tuning, and customer-facing technical engagements is highly desired.

Requirements

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, Information Technology, or related discipline.
  • 5+ years of experience designing and deploying high-performance networking solutions for AI infrastructure, HPC environments, GPU clusters, or Hyperscale cloud environments.
  • Experience designing and validating large-scale network fabrics supporting AI and distributed compute workloads.
  • Experience with customer-facing technical consulting, architecture reviews, or deployment leadership.

Nice To Haves

  • Expert-level knowledge of NVIDIA Spectrum-X, Spectrum-4 SN5600 Ethernet, Quantum InfiniBand (HDR/NDR/XDR), BlueField DPUs, DOCA SDK, RDMA and RoCEv2.
  • Experience supporting large-scale GPU cluster deployments and production AI environments.
  • Deep understanding of GPUDirect, NCCL, AI workload communication patterns, and GPU cluster validation and performance tuning.
  • Expertise in Adaptive Routing, SHARP, Congestion control, and Rail-optimized InfiniBand architectures.
  • Strong experience with BGP, EVPN, VXLAN, MPLS, OSPF, IS-IS, and Leaf-Spine architectures.
  • Experience designing lossless Ethernet environments utilizing PFC, ECN, DCQCN, and QoS.
  • Knowledge of high-performance storage networking and end-to-end infrastructure optimization.
  • Experience with Python, Ansible, Terraform, and REST APIs.
  • Experience implementing infrastructure observability, telemetry, monitoring, and troubleshooting solutions for large-scale networks.
  • Familiarity with data center operations, performance analytics, and capacity planning.
  • Preferred certifications include: NVIDIA Certified Networking Professional, NVIDIA InfiniBand Specialist, Cisco CCNP Data Center, Cisco CCIE Data Center, OCI Networking Specialist, AWS Solutions Architect Associate.

Responsibilities

  • Design GPU cluster network architectures for NVIDIA AI Factory environments utilizing Spectrum-X Ethernet (Spectrum-4 SN5600 + BlueField-3 DPUs) for enterprise deployments and Quantum XDR InfiniBand and Spectrum Ethernet for rack-scale GigaFactory deployments.
  • Design and validate large-scale AI fabrics supporting RDMA, RoCEv2, GPUDirect, NCCL collectives, and high-bandwidth GPU-to-GPU communications.
  • Develop rail-optimized InfiniBand topologies and Spectrum-X Adaptive Routing, SHARP, congestion management, and performance optimization strategies for AI training and inference environments.
  • Contribute to architecture decisions supporting large-scale distributed AI and HPC workloads.
  • Lead network bring-up, validation, and production-readiness activities for GPU infrastructure.
  • Develop and validate RDMA/RoCEv2 configuration guides and deployment runbooks for AI and HPC environments.
  • Establish performance baselines for RDMA throughput, RoCEv2 latency, GPU communication efficiency, storage and network throughput, and end-to-end infrastructure readiness.
  • Design failure-domain isolation strategies and resilient network architectures for large-scale AI deployments.
  • Support network architecture for high-density liquid-cooled GPU environments with power-aware design considerations.
  • Diagnose and resolve complex InfiniBand, Ethernet, RDMA, and GPU workload performance issues.
  • Analyze congestion, telemetry, traffic patterns, link utilization, routing behavior, and fabric health to identify bottlenecks and optimize performance.
  • Support root cause analysis and remediation of network, storage, and infrastructure issues impacting AI workload performance.
  • Develop operational runbooks and troubleshooting procedures consumed directly by Lenovo field teams and customers.
  • Design multi-tenant networking architectures supporting AI Factory, NeoCloud, and managed-service provider environments.
  • Implement namespace isolation, east-west traffic segmentation, secure tenant separation, and site resiliency aligned with NVIDIA Cloud Partner Reference Architectures.
  • Design BlueField-3 DPU solutions utilizing DOCA for infrastructure offload, security services, observability, and service mesh capabilities.
  • Collaborate with Lenovo engineering, product, ISG, and services organizations to validate designs against NVIDIA reference architectures and future hardware roadmaps.
  • Work directly with customers, partners, and delivery teams to translate AI workload requirements into production-ready infrastructure solutions.
  • Provide technical leadership during customer engagements, infrastructure deployments, escalations, and architecture reviews.
  • Support global deployments through approximately 40-50% travel, including customer workshops, implementation support, solution validation, and executive-level technical discussions.

Benefits

  • Hybrid Schedule on campus in Morrisville, NC. 3 days in office, 2 days work from home.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service