Lab Operations Engineer

TEKsystemsSanta Clara, CA
$70 - $75Onsite

About The Position

We are seeking a Lab Operations Engineer to support the bring-up, validation, diagnostics, and operational readiness of next-generation high-performance computing (HPC) and AI infrastructure. This role is highly hardware-focused and will play a critical part in evaluating system reliability, reproducing field issues, validating hardware platforms, and enabling internal engineering teams to successfully deploy workloads on on-prem systems. The ideal candidate combines hands-on lab experience with strong Linux systems administration skills and a passion for hardware diagnostics, validation, and system performance.

Requirements

  • Linux
  • Hardware Validation
  • System Validation
  • BIOS
  • Python
  • Bash
  • Automation
  • Experience with HPC, AI infrastructure, or large-scale compute clusters.
  • Background in validation engineering, system testing, or platform qualification.
  • Exposure to GPU-based systems and accelerated computing environments.
  • Experience with Dell, Supermicro, or other enterprise server platforms.
  • Knowledge of Linux kernel tuning, performance optimization, and hypervisor technologies.
  • Experience developing automated test frameworks and validation pipelines.
  • Understanding of datacenter infrastructure, rack-scale deployments, and operational workflows.
  • Familiarity with EDA, AI/ML, or other high-performance workloads.

Responsibilities

  • Investigate hardware failures and operational issues observed in production environments.
  • Reproduce large-scale system issues within a controlled lab environment.
  • Perform root cause analysis on server, memory, storage, networking, and thermal-related failures.
  • Execute component-level diagnostics, stress testing, and hardware validation procedures.
  • Analyze hardware fallout trends and collaborate with engineering teams to identify corrective actions.
  • Perform system bring-up activities for new server platforms and AI/HPC infrastructure.
  • Configure and validate hardware from OEM partners including Dell, Supermicro, and other server vendors.
  • Install, configure, and validate firmware, BIOS, BMC, and hardware management components.
  • Execute validation plans and test scripts for next-generation hardware platforms.
  • Support qualification efforts for new system architectures and deployment models.
  • Build and maintain recurring test pipelines for hardware validation and operational readiness.
  • Execute engineering design automation (EDA) and HPC workloads to verify system stability and performance.
  • Develop repeatable processes for system qualification and regression testing.
  • Partner with engineering teams to ensure hardware platforms meet operational requirements before deployment.
  • Administer Linux-based systems used for validation and testing.
  • Configure system-level settings including kernel parameters, drivers, and hypervisor configurations.
  • Support HPC cluster environments and distributed computing infrastructure.
  • Troubleshoot operating system and platform-level issues impacting performance or reliability.
  • Work closely with hardware engineering, datacenter operations, platform architecture, and vendor partners.
  • Collaborate with OEMs and suppliers to resolve hardware issues and improve platform quality.
  • Enable internal users and engineering teams to successfully deploy workloads on validated systems.
  • Support operational readiness efforts for emerging hardware technologies.

Benefits

  • Medical, dental & vision
  • Critical Illness, Accident, and Hospital
  • 401(k) Retirement Plan – Pre-tax and Roth post-tax contributions available
  • Life Insurance (Voluntary Life & AD&D for the employee and dependents)
  • Short and long-term disability
  • Health Spending Account (HSA)
  • Transportation benefits
  • Employee Assistance Program
  • Time Off/Leave (PTO, Vacation or Sick Leave)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service