High Performance Compute (HPC) Software Engineer – HPC SW Systems

KLA•Ann Arbor, MI
•$105,900 - $155,300•Onsite

About The Position

The Semiconductor Products and Customers (Semi PC) business unit designs, builds and sells KLA’s product portfolio to help chip manufacturers meet high-quality standards for technologies like AI, data centers, automotive and electronic devices. Our inspection, metrology, and analytics systems detect and monitor defects across logic and memory chips, specialty semiconductors, wafers, reticles, advanced packaging, process equipment and materials. We also develop technologies for specialty processes, IC substrates, and PCB manufacturing, including etch and deposition, component inspection, and imaging and analytics. KLA's Central Engineering organization is made up of nine Centers of Excellence (CoEs) spanning disciplines like automation, motion control, sensors, platform design, and packaging. Each CoE contributes not only technical deliverables but also deep expertise, best practices, and advanced tools that elevate both what we build—and how we build it. In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • Strong experience developing HPC or systems software on Linux.
  • Proficiency in Java and/or C++ and/or other system-level or performance-oriented languages.
  • Hands-on experience with parallel computing (MPI, OpenMP, multithreading).
  • Solid understanding of HPC hardware fundamentals: CPUs, memory hierarchies, storage, networking (Ethernet / InfiniBand).
  • Practical experience working with clusters, servers, or rack-scale systems in lab or production environments.
  • Strong debugging skills across software, OS, and hardware boundaries.

Nice To Haves

  • GPU computing (CUDA, ROCm, or equivalent) would be preferred.
  • Experience with containerized HPC environments (Docker, Singularity/Apptainer, Kubernetes in HPC contexts).
  • Familiarity with high-speed interconnects, storage architectures, and performance benchmarking.
  • Exposure to rack integration, including cabling, power distribution, cooling, and system bring-up.
  • Experience in semiconductor, manufacturing, or high-reliability systems environments.
  • Ability to reason about system reliability, MTBF/MTBA, and failure modes in large compute installations.

Responsibilities

  • Design, develop, and optimize HPC software running on large-scale Linux clusters, including distributed and parallel workloads (MPI, multithreading, GPU-accelerated pipelines, containerized workloads).
  • Optimize application performance and power utilization across CPU, memory, storage, and network subsystem, with attention to throughput, latency, and scaling behavior.
  • Develop and maintain system-level tooling for cluster bring-up, diagnostics, monitoring including component power usages, and health checks.
  • Work closely with algorithms, systems and application teams to understand and translate workload characteristics into power-efficient HPC software solutions.
  • Collaborate with hardware and systems teams to define HPC node, storage, and interconnect requirements based on software and algorithm needs.
  • Understand and influence CPU/GPU selection, memory sizing, PCIe layout, NUMA behavior, and network topology to ensure optimal software performance.
  • Participate in HW/SW co-debug activities, including performance bottlenecks, stability issues, and failure analysis.
  • Understand rack-level integration of HPC systems, focusing on power, cooling, cabling, networking, and physical layout considerations.
  • Understand data-center and lab constraints such as power budgets, thermal limits, network drops, and serviceability.
  • Contribute to best practices, and design reviews for new platforms and refresh cycles.
  • Act as a technical bridge between software, hardware, systems teams.
  • Provide clear technical documentation covering software and system architecture, deployment flows, performance assumptions.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service