Neo Cloud - Principal Hardware Systems Engineer

Blaze Talent•Seattle, WA
•Remote

About The Position

Neo Cloud is seeking a Principal Hardware Systems Engineer to design, deploy, and manage its AI hardware infrastructure. This senior individual-contributor role requires an experienced system architect who can bridge customer needs, system architecture, and production operations. The ideal candidate will translate future AI/ML customer requirements into scalable hardware designs and operational strategies for a next-generation AI neocloud.

Requirements

  • 10+ years of professional hardware engineering experience building and operating hardware at massive scale.
  • Direct experience with AI-focused hardware deployments across different hardware vendors, including NVIDIA GPU platforms.
  • Experience with NVIDIA BLackwell, GB200 or GB300 NVL72.
  • Experience with NVLink, NVSwitch, InfiniBand or Spectrum-X Ethernet, ConnectX networking, and BlueField DPUs.
  • Deep system level expertise.
  • Proven experience operating large-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
  • Strong systems programming skills.
  • Excellent written and verbal communication skills.

Responsibilities

  • Identify customer requirements by engaging directly with customers, solutions architects, and product teams to understand future hardware needs for AI/ML workloads.
  • Extrapolate from current usage patterns and industry trends to anticipate future requirements.
  • Partner with vendors and drive their roadmaps.
  • Design, implement, and operate AI cloud hardware.
  • Architect, evaluate, and deploy NVIDIA GPU infrastructure, including Blackwell-based B200, GB200, GB300, DGX and NVL rack-scale systems.
  • Design rack-scale infrastructure encompassing NVLink and NVSwitch fabrics, Infiniband or Spectrum-X Ethernet with RoCE for scale-out networking; power delivery, liquid cooling, cabling and serviceability.
  • Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions.
  • Architect for the specific demands of AI workloads focusing on performance, security, resiliency and cost.
  • Design and implement consistent and efficient hardware operation.
  • Take end-to-end ownership of hardware from design, supply chain planning, deployment, operation and ultimately deprecation.
  • Ensure security and compliance.
  • Utilize strong debugging and analytical skills to identify failures across a variety of electrical, thermal, optical, mechanical and software failures.
  • Set technical direction and best practices for the hardware organization; author and review design documents for significant architectural changes.
  • Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
  • Influence technical strategy across adjacent teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service