VP - AI Infrastructure Engineering

Designworks TalentBellevue, WA
Hybrid

About The Position

AI Infrastructure Engineering Leader Bellevue, WA Area | Hybrid | Senior Leadership A rapidly growing, well-funded technology company is seeking an AI Infrastructure Engineering Leader to lead the engineering organization responsible for transforming newly deployed data center hardware into reliable, production-ready compute infrastructure. This is a high-impact leadership opportunity for someone who combines deep technical expertise in AI/HPC infrastructure with strong engineering leadership. The ideal candidate understands both the physical infrastructure layer and the software automation required to operate large-scale GPU and compute environments efficiently. You’ll operate at the intersection of servers, GPUs, Linux, networking, Kubernetes, distributed systems, automation, and infrastructure software, helping establish the architecture, standards, tooling, and engineering practices required to deploy and operate infrastructure at scale.

Requirements

  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines.
  • 5+ years of engineering leadership experience, including managing and developing highly technical teams.
  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure.
  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation.
  • Hands-on understanding of Kubernetes, containers, and infrastructure-as-code.
  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production.
  • Ability to operate comfortably across both hardware and software organizations.
  • Strong communication and cross-functional leadership skills.
  • A hands-on, high-ownership leadership style with the ability to operate effectively in a fast-moving, build-from-the-ground-up environment.

Nice To Haves

  • Experience in one or more of the following areas is highly valued:
  • GPU infrastructure, NVIDIA platforms, AI or HPC environments.
  • Bare-metal provisioning technologies such as MAAS, Ironic, xCAT, Foreman, or similar.
  • Hardware management technologies including Redfish, IPMI, BMC, PXE, and firmware management.
  • NVIDIA technologies such as CUDA, NVML, DCGM, NVIDIA drivers, or GPU Operator.
  • High-performance networking including InfiniBand, RoCE, RDMA, or high-speed Ethernet.
  • Cluster orchestration and scheduling technologies such as Kubernetes, Slurm, or similar.
  • Automated infrastructure validation and hardware health testing.
  • Experience scaling infrastructure across thousands of servers or GPUs.

Responsibilities

  • Lead the engineering organization responsible for data center infrastructure bring-up and production readiness.
  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness.
  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation.
  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure.
  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration.
  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies.
  • Oversee integration with technologies such as Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms.
  • Partner closely with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams to deliver production-ready infrastructure.
  • Establish automated testing, health checks, monitoring, and validation processes to identify infrastructure issues before workloads reach production.
  • Improve deployment speed, reliability, automation, scalability, and operational efficiency.
  • Build and develop a highly capable engineering organization while establishing processes that can scale with the business.

Benefits

  • The company operates with a startup mentality—fast, lean, highly collaborative, and high ownership—while having the resources to build infrastructure for significant scale.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service