Senior HPC Hardware Engineer

NorthMark StrategiesDallas, TX
Onsite

About The Position

NorthMark Compute & Cloud (NMC²) is seeking a Senior HPC Hardware Engineer to join the Compute Engineering team based at our Dallas, TX offices at Victory Commons. This is a hands-on role at the center of one of the most demanding and rapidly scaling HPC environments in operation, spanning a large fleet of GPU and CPU nodes built on the latest NVIDIA platforms including H200 and GB200/NVL72 architectures. You will own the full hardware lifecycle for NMC²’s compute fleet — from bare-metal provisioning and firmware baseline management through production validation, troubleshooting, and capacity planning. You will be the go-to expert on server hardware architecture, driving standards and automation that ensure the fleet operates at peak performance and availability. Your work directly enables the research and delivery workloads that NMC²’s clients depend on. This role requires close collaboration with Software Engineering, Networking, and Vendor teams, and involves mentoring junior engineers. The ideal candidate is a technically deep infrastructure leader who thrives in fast-paced environments, brings a strong automation mindset, and has a proven track record managing large-scale HPC or AI compute infrastructure.

Requirements

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands-on experience.
  • 5+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
  • Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.
  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
  • Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.
  • Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
  • Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.
  • Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Nice To Haves

  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.

Responsibilities

  • Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across NMC²’s Dallas infrastructure.
  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
  • Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.
  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
  • Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.

Benefits

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub
  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability
  • 401(k): Company will match 100% of your contributions up to 6%
  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.
  • Time Off: 25 days of Paid Time Off plus 12 company holidays
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service