Senior AI GPU Deployment Engineer

5C Data Centers USA, Inc.Springfield, OH
$120,000 - $150,000Remote

About The Position

We are seeking an experienced Senior AI GPU Deployment Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers. The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills.

Requirements

  • Bachelor's degree in Computer Science, Engineering, IT, or related field (or equivalent experience)
  • 5+ years of infrastructure engineering or datacenter deployment experience
  • 3+ years deploying large-scale AI, HPC, or GPU infrastructure
  • Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
  • Strong expertise with GPU architectures
  • Strong expertise with InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
  • Strong expertise with NVLink, NVSwitch, and GPU-direct technologies
  • Strong expertise with Canonical MaaS and automated provisioning systems
  • Strong expertise with VAST Data or similar high-performance storage platforms
  • Strong expertise with Linux systems administration for HPC/AI workloads
  • Strong expertise with Infrastructure-as-Code and configuration management (Ansible)
  • Strong expertise with Python, Shell, and SQL for infrastructure automation and diagnostics
  • Strong understanding of RDMA, RoCE, and lossless Ethernet fabrics
  • Strong understanding of Cluster automation, observability, and lifecycle management

Responsibilities

  • Deploy, integrate, and validate multi-rack GPU-based compute platform deployments.
  • Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM) and validate interconnect and fabric performance (nccl).
  • Collaborate with the network engineering team on topology implementation and optimization and the storage engineering team on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data).
  • Configure settings and manage firmware updates for GPUs, NICs, BMC, BIOS and other components across large-scale clusters.
  • Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management.
  • Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards.
  • Query and analyze deployment outcomes using SQL for diagnostics and operational reporting.

Benefits

  • Competitive pay
  • Meaningful, lasting impact on the work you do
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service