Senior AI GPU Deployment Engineer

5C Data Centers USA, Inc.•Springfield, OH
•$120,000 - $150,000•Remote

About The Position

We are seeking an experienced Senior AI GPU Infrastructure Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers. The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills.

Requirements

  • Bachelor's degree in Computer Science, Engineering, IT, or related field (or equivalent experience)
  • 5+ years of infrastructure engineering or datacenter deployment experience
  • 3+ years deploying large-scale AI, HPC, or GPU infrastructure
  • Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
  • Strong expertise with: GPU architectures
  • InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
  • NVLink, NVSwitch, and GPU-direct technologies
  • Canonical MaaS and automated provisioning systems
  • VAST Data or similar high-performance storage platforms
  • Linux systems administration for HPC/AI workloads
  • Infrastructure-as-Code and configuration management (Ansible)
  • Python, Shell, and SQL for infrastructure automation and diagnostics
  • Strong understanding of: RDMA, RoCE, and lossless Ethernet fabrics
  • Cluster automation, observability, and lifecycle management

Responsibilities

  • Deploy and integrate GPU-based compute platforms from NVIDIA and other accelerator vendors
  • Execute end-to-end deployment of multi-rack AI clusters in hyperscale datacenters
  • Support rack-and-stack, NVLink/NVSwitch cabling, fabric deployment, burn-in, and cluster validation
  • Validate deployment readiness, GPU fabric performance, acceptance testing, and operational handoff
  • Validate high-performance GPU interconnects based on InfiniBand and Ethernet GPU fabric architectures
  • Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM) and validate fabric performance (nccl)
  • Collaborate with network engineering teams on topology implementation and optimization
  • Coordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data)
  • Validate storage throughput, latency, and GPU data delivery performance
  • Manage firmware updates for GPUs, NICs, BMC, and other components across large-scale clusters
  • Configure and validate BIOS settings for HPC/AI workloads (CPU affinity, NUMA, C-states, power management)
  • Configure and manage BMC for remote infrastructure management (IPMI, Redfish)
  • Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management
  • Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards
  • Query and analyze deployment outcomes using SQL for diagnostics and operational reporting

Benefits

  • Competitive pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service