About The Position

NVIDIA is seeking experienced candidates with a passion for infrastructure security, compute services, and storage needs for manufacturing workloads. This role involves redesigning core services for NVIDIA manufacturing sites, enabling new sites, and troubleshooting incidents with a focus on minimizing revenue impact from factory downtime. It's a hands-on engineering position focused on developing and deploying ultra-high-speed, resilient, and scalable core services for GPU-accelerated OT and IT environments in manufacturing sites. Success requires outstanding problem-solving abilities, a comprehensive understanding of compute, storage, routing, switching, automation, and fundamental network theory.

Requirements

  • MS or PhD in Electrical Engineering, Computer Science, Computer Engineering, Artificial Intelligence, Data Science, Mathematics, Statistics, or equivalent experience.
  • 12+ years of experience in building, managing and supporting large scale hybrid networks.
  • Developing automation pipelines with Python, Ruby, Go or other languages used in infrastructure automation.
  • Expert in networking technologies especially Mellanox.
  • Expert in Compute technologies - Dell, Openshift, cloud compute services like Dell etc.
  • Expert in Storage Services - Pure, Netapp.
  • Experience designing compute and storage architectures for data centers, offices, Manufacturing environments and Labs.
  • Experience building data lakes, caching layers and cloud based infrastructure services with the ability to take requirements into new designs as businesses transform and product requirements change.
  • Experience with designing Test Automation infrastructures in manufacturing.
  • Strong scripting skills.
  • Designing both CPU and GPU workloads.

Responsibilities

  • Lead the architecture, design, and deployment of global-scale manufacturing sites and their connectivity to AI factories and offices.
  • Architect and build CPU-based compute, storage, and GPU/HPC clusters.
  • Design high-performance OT and IT networks for NVIDIA and partner connectivities to support general compute workloads and GPU-dense AI/ML training and inference environments.
  • Partner with systems, Operations teams, supply chain partners, OS, GPU, storage, and HPC product teams to deliver scalable, highly available network architectures and connectivity solutions.
  • Implement and refine compute, storage, security, telemetry, and performance-engineering practices across the infrastructure.
  • Manage life cycle management and design infrastructures for revenue-generating manufacturing sites.
  • Define and enforce security, compliance, and reliability standards for all infrastructure components supporting mission-critical manufacturing and R&D workloads.
  • Collaborate with Operations, Manufacturing Partners, and engineering teams to develop “NVIDIA on NVIDIA” reference architectures and best-practice solutions for large-scale compute and AI data designs.
  • Troubleshoot DNS, DHCP, and other full connectivity stack infrastructures.
  • Experience with Test Engineering topologies and supporting infrastructure on the factory floor.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service