Staff HPC Infrastructure Engineer

Guardant HealthPalo Alto, CA
Hybrid

About The Position

Guardant Health's High-Performance Computing (HPC) team is seeking a Staff-level engineer with broad HPC competency and specialized depth in areas such as Red Hat-family OS management, networking, storage, Kubernetes, or Slurm. This role involves building and operating the company's computational technology backbone, including scalable data storage, high-performance compute clusters, and software infrastructure. The engineer will contribute to both on-premise and cloud environments, partnering with internal teams and managed service providers to scale operations and maintain infrastructure. The position requires strong technical skills, the ability to manage multiple projects, and a commitment to engineering excellence in a fast-paced, agile environment. The role also involves maintaining day-to-day support SLAs while driving key projects forward and acting as a deep technical expert in their specialization.

Requirements

  • Bachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experience
  • Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking.
  • Hands-on experience with automation tools, such as Ansible or equivalent technologies.
  • Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments.
  • Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments.
  • Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments.
  • Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets.
  • Strong experience creating and maintaining system administration and technical documentation.
  • Red Hat family OS management is a must.

Nice To Haves

  • Cisco Certified Network Professional (CCNP) certification
  • Experience with Arista and compatible networking, up to and including 400 Gb/s links
  • Experience administering IBM's General Parallel File System (GPFS)
  • Experience administering the Slurm scheduler
  • Experience using Warewulf Linux support and OS management.
  • Debian or Suse experience
  • Experience with cloud bursting technologies
  • Experience with wide area file systems
  • Experience with Docker and Apptainer container technologies
  • Experience with Kubernetes
  • Operating infrastructure compliant with HIPAA and SOX standards

Responsibilities

  • Manage multiple HPC clusters and cluster file systems
  • Integrate cloud bursting as part of the HPC abstraction work
  • Research, develop, and implement the next generation HPC solutions
  • Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others
  • Maintain, monitor, and support the infrastructure environment and/or facilities
  • Use and maintain enhanced production monitoring and addition capability
  • Support improvements for increased system reliability and performance
  • Support multiple systems or applications of medium to high complexity with multiple concurrent users, ensuring control, integrity, and accessibility
  • Support systems at remote locations, including internationally
  • Mentor junior engineers on HPC best practices
  • Work with offsite consultants to maintain the infrastructure
  • Work with vendors to troubleshoot, upgrade, and repair systems as needed
  • Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP
  • Set up and ownership supporting XDMoD instances for HPC metric and monitoring
  • Participate in a 24/7 on-call rotation
  • Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer
  • Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet
  • Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure
  • Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locations
  • Act as the technical peer for the architecture and integration strategy for HPC storage in partnership with the dedicated storage engineer and MSP
  • Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and escalate technical issues
  • Support the transition of day-to-day storage operations to the MSP without loss of performance or reliability

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • Disability insurance
  • 401k
  • Paid holidays
  • Flexible scheduling
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service