Hardware Diagnostics Engineer - Infrastructure

TensorWave•Las Vegas, NV
•Onsite

About The Position

TensorWave is seeking a Hardware Diagnostics Engineer to join their team. This role involves running burn-in tests, triaging hardware failures, managing out-of-band server operations, and overseeing the entire RMA (Return Merchandise Authorization) process with vendors. The engineer will be responsible for ensuring that all GPU servers are thoroughly tested and proven to be functional under load and temperature before being deployed for customer workloads. This includes diagnosing issues, coordinating hardware replacements, and managing the return of faulty components.

Requirements

  • 3–6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering
  • Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, and reading hardware behavior at the rack
  • Practical experience with BMCs and out-of-band management: IPMI, Redfish, iDRAC, iLO, or equivalent
  • Strong Linux troubleshooting: boot process, driver and device issues, and diagnostic tools such as dmesg, lspci, ipmitool, and SMART
  • Comfort reading sensor data, event logs, and thermal and power telemetry well enough to tell a real failure from noise
  • Working scripting ability in Bash or Python — enough to automate a repetitive task and read someone else's tooling
  • Experience running hardware RMAs with vendors, or a clear track record of driving issues to closure with outside parties
  • A methodical troubleshooting habit: you isolate variables, you don't change three things at once, and you can say what evidence led to your conclusion
  • Clear written communication for tickets, runbooks, and vendor cases

Nice To Haves

  • GPU server experience, especially AMD GPUs and ROCm
  • Burn-in, stress testing, or node validation tooling in a GPU or HPC environment
  • Familiarity with firmware update processes and why fleet-wide changes get staged
  • NetBox or other DCIM and IPAM tooling
  • Ansible, or Python against REST APIs
  • Prior work in a high-volume hardware environment: hyperscaler, colo, integrator, or manufacturing test

Responsibilities

  • Run server and GPU burn-in and stress testing, interpret the results, and decide whether hardware is production-ready
  • Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling: reproduce the failure, isolate the faulty component, and document what proved it
  • Work servers out-of-band through IPMI and Redfish for power control, boot configuration, BIOS settings, and sensor and event log collection
  • Apply firmware updates across the fleet following the team's qualified baselines and rollout process
  • Drive RMAs with vendors from ticket through replacement, installation, and return of the failed part
  • Keep asset, serial, and replacement history accurate in NetBox so we know what's actually in every rack
  • Track failure patterns across the fleet and raise them when the same part or firmware version keeps turning up
  • Improve the runbooks you work from, and script the steps you find yourself repeating
  • Partner with datacenter operations on hands-on work during turn-ups and expansions
  • Take part in an on-call and escalation rotation for hardware issues

Benefits

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service