Sr. System Engineer/GPU Platforms

SupermicroSan Jose, CA

About The Position

Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification, enablement, and customer deployment of advanced GPU computing platforms. This role focuses on multi-GPU server systems used for AI, HPC, enterprise computing, and accelerated workloads. The successful candidate will work across the product lifecycle, from initial system bring-up and qualification through product release, customer POC/EVAL support, debugging, and post-launch technical enablement. The ideal candidate combines strong server hardware knowledge with hands-on Linux and GPU software experience and can independently troubleshoot complex issues across hardware, firmware, operating systems, networking, and GPU software environments.

Requirements

  • Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience.
  • 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field.
  • Strong knowledge of enterprise server hardware and system architecture.
  • Hands-on experience with Linux server environments.
  • Experience installing, configuring, validating, and troubleshooting server hardware and software.
  • Strong system-level troubleshooting and root-cause-analysis skills.
  • Working knowledge of PCIe architectures and high-performance I/O.
  • Experience with GPU computing, accelerators, or comparable high-performance computing technologies.
  • Ability to independently manage complex technical assignments and drive issues toward resolution.
  • Strong written and verbal communication skills.
  • Ability to work effectively with cross-functional and geographically distributed engineering teams.
  • Comfortable participating in customer-facing technical discussions.

Nice To Haves

  • Hands-on experience with NVIDIA data center or professional GPU platforms.
  • Experience with CUDA and NVIDIA GPU software environments.
  • Experience with NVIDIA NVQUAL or similar platform qualification processes.
  • Experience with 4-GPU or 8-GPU server platforms.
  • Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures.
  • Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication.
  • Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet.
  • Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools.
  • Experience with Docker, containers, Kubernetes, or related orchestration technologies.
  • Experience supporting AI, machine learning, HPC, or accelerated computing environments.
  • Experience with customer POCs, technical evaluations, or engineering escalations.
  • Experience delivering technical training or knowledge-sharing sessions.
  • Bash, Python, or other scripting experience is a plus.

Responsibilities

  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes.
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components.
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation.
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities.
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution.
  • Collaborate with internal Architecture, Systems, Software, Validation, Product Management, and other engineering teams, as well as external technology partners.
  • Develop technical documentation, troubleshooting guides, and best practices.
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities.
  • Serve as a technical resource and mentor for other engineers when appropriate.

Benefits

  • comprehensive benefits package
  • participation in bonus and equity award programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service