High Performance Computing (HPC) - Software Engineer

KLAMilpitas, CA
$94,600 - $160,800Onsite

About The Position

Develop, deploy, and support software infrastructure for High-Performance Computing (HPC) systems used for semiconductor data processing, computational lithography, AI workloads, and distributed compute platforms. Work with Linux-based clusters, schedulers, storage systems, networking fabrics, virtualization platforms, and automation frameworks to improve reliability, scalability, and performance.

Requirements

  • Bachelor's Level Degree and 0-2 years related work experience
  • BS/MS in Computer Engineering, Computer Science, Electrical Engineering, or related field.
  • 2+ years of software development or systems engineering experience.
  • Strong Linux fundamentals.
  • Experience with Python and Shell scripting.
  • Experience with Git and modern software development practices.
  • Understanding of TCP/IP networking, DNS, DHCP, NFS, and system services.
  • Familiarity with virtualization technologies such as VMware, Proxmox, or KVM.

Nice To Haves

  • HPC cluster experience.
  • SGE, SLURM, Kubernetes, or container technologies.
  • GPU computing (CUDA/NVIDIA ecosystem).
  • Configuration management (Ansible, Puppet, Chef, Salt).
  • CI/CD frameworks and DevOps workflows.
  • Storage technologies including NFS, NAS, RAID, and parallel file systems.
  • Independently implement features and fixes.
  • Own small-to-medium technical projects.
  • Troubleshoot and resolve operational issues.
  • Contribute to architecture discussions.

Responsibilities

  • Develop and maintain Python, Bash, and system automation tools.
  • Support Linux-based HPC clusters (RHEL, Rocky, AlmaLinux, Ubuntu).
  • Assist with deployment and maintenance of compute nodes, storage, and networking infrastructure.
  • Configure and troubleshoot schedulers such as SGE, SLURM, or equivalent.
  • Develop tooling for cluster monitoring, provisioning, and software deployment.
  • Collaborate with software, algorithm, systems, manufacturing, and field teams.
  • Debug performance, scalability, and reliability issues across distributed systems.
  • Participate in code reviews, CI/CD, testing, and release activities.
  • Create operational runbooks, deployment procedures, and technical documentation.
  • Support customer and manufacturing escalations.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service