HPC & GPU Infrastructure Support Engineer

Vast.aiLos Angeles, CA
$90,000 - $150,000Onsite

About The Position

This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM-based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack. You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end-to-end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams. The best engineers in this role don’t just resolve individual issues—they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems. Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads. Vast.ai users or hosts strongly preferred.

Requirements

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions
  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting
  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting
  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting
  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting
  • Python and Bash scripting skills for automation and diagnostic tooling
  • Strong written English communication that is clear, professional, and technically precise
  • Experience providing technical support in a customer-facing or internal help desk environment
  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice To Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers
  • Monitoring and observability experience (Prometheus, Grafana)
  • Relevant certifications: RHCSA, CompTIA Linux+, or similar
  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Responsibilities

  • Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments
  • Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks
  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads
  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines
  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration
  • Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post-setup troubleshooting
  • Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance
  • Provide coverage for L1 support overflow during peak periods or incidents
  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations
  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead
  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

Benefits

  • Comprehensive health, dental, vision, and life insurance
  • 401(k) with company match
  • Meaningful early-stage equity
  • Onsite meals, snacks, and close collaboration with founders/tech leaders
  • Ambitious, fast-paced startup culture where initiative is rewarded
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service