Engineering Manager, GPU Infrastructure

5C Data Centers USA, Inc.Springfield, OH
$180,000 - $220,000Remote

About The Position

We are seeking an experienced Engineering Manager, GPU Infrastructure to lead the planning, deployment, integration, and operational readiness of large-scale AI infrastructure environments. This role is responsible for delivering production-grade GPU clusters that support AI training, inference, and high-performance computing workloads across cloud, hybrid, and on-premises environments. The ideal candidate brings deep technical expertise in GPU infrastructure, networking, storage, automation, and datacenter deployment, combined with strong program leadership and cross-functional execution skills. This leader will oversee end-to-end AI cluster deployment initiatives, including hardware integration, rack-and-stack operations, provisioning automation, performance validation, and operational handoff. The role requires hands-on familiarity with modern AI infrastructure tooling and architectures, including Canonical MaaS, VAST Data storage platforms, and both InfiniBand and Ethernet-based GPU networking fabrics.

Requirements

  • Bachelor's degree in Computer Science, IT, or related field, plus 10+ years of infrastructure engineering or datacenter deployment experience.
  • 5+ years leading deployment or operations teams focused on large-scale AI, HPC, or GPU infrastructure, with a proven track record managing complex cross-functional programs.
  • Hands-on experience deploying and operating large enterprise/hyperscale GPU clusters, alongside expert Linux systems administration skills.
  • Proficiency with Canonical MaaS, data storage platforms, network architecture, and GPU server architectures.
  • In-depth knowledge of InfiniBand, Ethernet GPU fabrics, RDMA/RoCE networking, high-performance storage, and cluster provisioning automation.

Responsibilities

  • Lead the deployment, logical integration, and validation of large-scale NVIDIA/accelerator GPU clusters, including rack-and-stack, burn-in testing, and operational turnover.
  • Oversee high-performance GPU interconnects (InfiniBand/Ethernet fabrics), optimizing spine-leaf architectures, RDMA, network telemetry, and overall performance.
  • Partner with storage teams to integrate, tune, and validate data platforms to guarantee high-throughput, low-latency data delivery for AI workloads.
  • Drive cluster provisioning using Infrastructure-as-Code, Canonical MaaS, and observability tools to establish repeatable, automated deployment standards.
  • Partner with PMO, datacenter ops, and vendor teams to manage full-stack deployment programs while tracking milestones in systems like Jira.
  • Build and mentor high-performing engineering teams while creating operational documentation, standardized best practices, and scalable deployment frameworks.

Benefits

  • Competitive pay plus meaningful, lasting impact on the work you do.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service