Senior HPC Software & Systems Architect

KLAMilpitas, CA
$186,200 - $316,500Onsite

About The Position

In this exciting role you will design and architect next-generation High Performance Computing platforms supporting semiconductor manufacturing, AI workloads, computational lithography, metrology, inspection, and large-scale distributed data processing environments. Lead architecture decisions spanning software, hardware, networking, storage, GPU acceleration, virtualization, and cloud-integrated HPC infrastructure. Drive technology strategy, roadmap planning, and cross-functional alignment from concept through field deployment.

Requirements

  • Expert Linux knowledge: RHEL, Rocky, AlmaLinux, Ubuntu, SUSE
  • HPC Technologies: SGE, SLURM, Kubernetes, Distributed computing architectures, NUMA optimization, MPI concepts, GPU workload scheduling
  • Infrastructure: NVIDIA GPU ecosystems, High-speed Ethernet, RDMA, NVMe storage, Parallel file systems, NAS/NFS architectures
  • Virtualization: VMware, Proxmox, KVM
  • Container technologies
  • Software Development: Python, C++, Bash, Git, CI/CD pipelines, Ansible, Infrastructure automation frameworks
  • System Architecture: Service-oriented design, Scalability engineering, Reliability engineering, High availability, Disaster recovery, Capacity modeling, Performance benchmarking and tuning
  • Doctorate (Academic) Degree and related work experience of 5 years; Master's Level Degree and related work experience of 8 years; Bachelor's Level Degree and related work experience of 12-15 years

Responsibilities

  • Define architecture for enterprise-scale HPC clusters.
  • Develop compute, network, storage, and virtualization strategies.
  • Lead platform scalability planning from tens to thousands of cores.
  • Define standards for GPU acceleration, distributed computing, and workload orchestration.
  • Establish system performance entitlement targets and capacity planning methodologies.
  • Design distributed software services supporting HPC infrastructure.
  • Define APIs, automation frameworks, configuration management architecture, and observability solutions.
  • Lead adoption of software engineering best practices including CI/CD, testing, infrastructure-as-code, and DevOps methodologies.
  • Drive modernization initiatives across Linux, container, and cloud-native platforms.
  • Provide technical leadership across software, systems, manufacturing, operations, and field organizations.
  • Lead architecture reviews and design reviews.
  • Mentor engineers across multiple levels.
  • Establish coding, deployment, and operational standards.
  • Serve as escalation leader for critical system issues affecting customers and manufacturing operations.
  • Define HPC platform roadmap.
  • Evaluate emerging technologies in: GPU computing, AI infrastructure, Storage architectures, Virtualization, Cloud integration, High-speed networking
  • Partner with product management and business leadership on long-term HPC strategy.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume

What This Job Offers

Job Type

Full-time

Career Level

Senior

Education Level

Ph.D. or professional degree

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service