Network Engineer, Platform, Automation & HPC/AI

Lawrence Berkeley National LaboratoryBerkeley, CA
$139,440 - $267,996Hybrid

About The Position

Join NERSC at Berkeley Lab and help engineer the high-performance network platform powering some of the nation’s most advanced scientific computing. As a Network Engineer, Platform, Automation & HPC/AI, you’ll work at the intersection of network engineering, automation, and software development while helping advance AI-driven network operations. Our team manages 1 Tb/s of border connectivity to ESnet and an 800G/400G data center network backbone supporting the NERSC-9 and NERSC-10 supercomputers, multi-tier storage, archive, and edge services. Your work will help improve the performance, scalability, automation, and reliability of scientific workflows serving more than 10,000 users. Our mission is to bring science solutions to the world. We welcome candidates from all backgrounds, including those with non-traditional paths. We value a growth mindset and believe skills and experience are transferable. If you’re eager to learn and meet the qualifications below, we encourage you to apply. Join our team where your work can have a high impact for an organization associated with 17 Nobel Prizes… and counting. This position may be filled at Level 3 or Level 4. Level 3 is intended for experienced engineers who independently solve complex networking and automation challenges. Level 4 is intended for senior technical leaders who architect solutions, lead modernization efforts, and establish new technical approaches for complex HPC and data center environments.

Requirements

  • Typically requires a minimum of 8 years of related experience with a Bachelor’s degree; or 6 years and a Master’s degree; or equivalent experience.
  • 5+ years of IP networking experience in Data Center or LAN environments.
  • 2+ years of experience with network automation and/or network performance optimization.
  • Project management experience in developing technical project scope, schedule and budget.
  • Strong written, verbal, listening, and presentation skills; effective collaboration skills with technical peers, vendors, and customers.
  • Ability to resolve complex issues in creative and effective ways.
  • Ability to network and collaborate with key contacts outside their own area of expertise.
  • Demonstrated ability to work effectively as part of a cross-disciplinary team.
  • Typically requires a minimum of 12 years of related experience with a Bachelor’s degree; or 8 years and a Master’s degree; or equivalent experience.
  • 10+ years of IP networking experience in Data Center or LAN environments.
  • 4+ years of experience with network automation and/or network performance optimization.
  • Demonstrated project management expertise, specifically, in leading the development of technical project scope, schedule and budget.
  • Excellent written, verbal, listening, and presentation skills; excellent communication and collaboration skills with technical peers, vendors, and customers.
  • Expert-level capabilities in configuring, troubleshooting, and using IPv4/IPv6 routing protocols, preferably BGP, VXLAN, L3VPN, OSPF, and/or ISIS in a WAN or LAN environment.
  • Strong software development skills in Python or Go and experience with network automation tools (Ansible, Terraform, eAPI, etc.).
  • Demonstrated experience with AI-driven networking, intent-based networking, or ML-based network optimization.
  • Demonstrated experience in leading network technology migrations or upgrades in high-performance computing environments.
  • Demonstrated experience in automating the network’s daily operational, deployment and monitoring tasks.
  • Demonstrated experience with RDMA (Remote Direct Memory Access), InfiniBand, HPC Protocol, RoCE (RDMA over Converged Ethernet).
  • Demonstrated experience working with Network Management platforms like Arista CloudVision, NVIDIA UFM, Dell SFM.
  • Demonstrated experience in designing and implementing network security architectures, including segmentation (e.g., VRFs, micro-segmentation), access control (ACLs), encryption (IPsec/MACsec), and threat detection/mitigation in large-scale enterprise or HPC environments.
  • Demonstrated experience in hybrid cloud networking, including connectivity models (VPN, Direct Connect/ExpressRoute), multi-cloud architectures, and integrating on-prem HPC or data center networks with public cloud environments.

Responsibilities

  • Implement, operate, maintain, and improve network automation and observability solutions.
  • Contribute to Data Center modernization efforts and NERSC’s Smart Facility initiative.
  • Support efforts to design and deliver network services to address emerging needs (e.g., American Science Cloud, new Edge services).
  • Continuously monitor and optimize network performance, focusing on reducing latency, maximizing throughput, and improving fault tolerance.
  • Create and maintain comprehensive network documentation, including physical and logical topology diagrams.
  • Collaborate with the Security Group to implement security measures for data integrity and privacy, ensuring high availability and reliability through redundancy and failover mechanisms.
  • Share on-call rotation with colleagues and serve as an escalation contact for service incidents.
  • Work on and resolve complex issues where analysis of situations or data requires an in-depth evaluation of multiple variables.
  • Exercise judgment in selecting methods, techniques and evaluation criteria for obtaining results.
  • Build effective working relationships with technical partners and stakeholders across disciplines.
  • Architect, develop, and establish technical direction for network automation, observability, and self-healing capabilities.
  • Lead Data Center modernization efforts in support of NERSC’s Smart Facility initiative and emerging needs e.g., American Science Cloud
  • Design, develop, and maintain automation frameworks, infrastructure-as-code, and software solutions to manage, optimize, and self-heal the HPC and data center network.
  • Build and integrate AI/ML-driven observability, predictive analytics, and automated remediation capabilities into network operations.
  • Independently develop technical approaches to complex and novel networking problems.
  • Exercise significant technical judgment when evaluating architecture, tools, and implementation approaches.
  • Establish methods and technical practices for new or ambiguous assignments.

Benefits

  • Exceptional health and retirement benefits, including pension or 401K-style plans
  • Opportunities to grow in your career - check out our Tuition Assistance Program
  • A culture where you’ll belong - we are invested in our teams!
  • In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year.
  • Parental bonding leave (for both mothers and fathers)
  • Pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service