Senior GPU Engineer

Vultr
•$190,000 - $210,000•Remote

About The Position

Vultr is seeking a highly skilled and experienced Senior GPU Engineer to design, optimize, and scale GPU infrastructure for AI training and inference workloads. The ideal candidate has deep experience with GPU systems, distributed performance tuning, and leading engineering initiatives across cross-functional teams. This is a highly visible role in a high-growth technology company, which will require ownership of end-to-end system performance, from hardware qualification to cluster-level optimization, and the ability to mentor and elevate team capabilities. This is your opportunity to join our fast growing team and leave your mark on Vultr and the future of Cloud Infrastructure.

Requirements

  • 5+ years of experience in GPU infrastructure, HPC, or distributed systems
  • Strong expertise in Linux systems and server hardware
  • Proven experience with large-scale GPU clusters
  • Strong programming skills in Python (beyond basic scripting)
  • Experience with automation frameworks (Ansible or similar)
  • Experience defining validation standards and performance baselines for GPU infrastructure
  • Strong debugging skills across system layers (hardware → OS → network)
  • Familiarity with high-speed networking concepts to coordinate with fabric engineering teams
  • Excellent communication and cross-team collaboration skills

Responsibilities

  • Own end-to-end validation of GPU clusters and new hardware platforms
  • Lead hardware qualification and bring-up for new GPU platforms
  • Analyze and resolve performance bottlenecks across GPU, CPU, PCIe, and network layers
  • Establish and maintain validation frameworks, test suites, and performance baselines
  • Develop and enhance automation frameworks for cluster provisioning and validation
  • Define performance baselines and validation methodologies
  • Troubleshoot complex distributed system issues, including communication libraries (e.g., NCCL)
  • Improve system reliability through proactive testing and tuning
  • Mentor engineers and elevate team capabilities
  • Drive cross-functional initiatives to improve GPU cluster reliability and efficiency

Benefits

  • 100% company-paid insurance premiums for employee medical, dental and vision plans
  • 401(k) plan that matches 100% up to 4%, with immediate vesting
  • Professional Development Reimbursement of $2,500 each year
  • 11 Holidays + Paid Time Off Accrual + Rollover Plan
  • Increased PTO at 3 year and 10 year anniversary
  • 1 month paid sabbatical every 5 years
  • Anniversary Bonus each year
  • $500 stipend for remote office setup in first year + $400 each following year
  • Internet reimbursement up to $75 per month
  • Gym membership reimbursement up to $50 per month
  • Company paid Wellable subscription
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service