High-Performance Computing (HPC) Engineer

The Aerospace CorporationEl Segundo, CA
$135,200 - $202,800Onsite

About The Position

The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative, proactive mindset and a shared commitment to engineering excellence. The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • Minimum of 7 years’ experience in Linux system administration within an enterprise HPC environment.
  • Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules.
  • In-depth knowledge of Linux, networking, and HPC systems.
  • Experience with Infrastructure-as-Code and GitOps
  • Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
  • Experience provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA)
  • Proficiency in scripting and competence with automation tools such as Clush.
  • Experience hardening Linux systems to meet security requirements
  • Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.
  • Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.
  • CompTIA Security+ CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel
  • Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required).
  • Demonstrated ability to lead cross-functional teams and mentor junior engineers.
  • 9+ years of experience in an enterprise 100+ server HPC cluster operations and administration
  • Experience with performance analysis and optimization with custom developed technical software in collaboration with scientists and engineers.
  • Expertise in optimizing and customizing Slurm partitions, qualities of service and priority to balance utilization and reduce job wait times.
  • Experience performing in-place upgrades of Slurm.
  • Implementing visualization of live system telemetry
  • Experience developing and architecting cluster configuration management
  • Advanced Infrastructure-as-Code GitOps (e.g. multi-branch pipelines)
  • Skill in provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA, Nsight), with expertise in GPU integration, resource allocation, and scheduling using Slurm.

Nice To Haves

  • An active TS/SCI clearance with CI Polygraph.
  • Experience integrating Slurm with SELinux
  • Experience implementing DISA STIG compliance
  • Experience integrating Kubernetes and Slurm, e.g. Slinky
  • Experience supporting and managing diverse HPC workloads, including computational fluid dynamics, Monte Carlo, structural analysis
  • Experience integrating user web portals to launch & manage workloads (e.g. Open OnDemand and developing plugins for session persistence and VSCode)
  • Experience implementing utilization dashboards (e.g. XDMod) with Slurm
  • Experience managing parallel file systems such as Lustre.
  • Experience developing solutions that optimize data storage
  • Experience implementing or supporting Slurm REST API
  • Experience with automated provisioning, e.g. Warewolf, Kickstart, PXE
  • Knowledge of NVLINK and DCGM for optimizing GPU workflows.
  • Familiarity with Prometheus and Grafana for monitoring and performance visualization.
  • Background in containerization within an HPC context used for data processing and technical analysis
  • Experience packaging custom software (e.g. RPMs)
  • Proficiency with automation tools such as Ansible for HPC
  • Hands-on background with cloud HPC services
  • Experience with AWS Parallel Computing Service (AWS ParallelCluster).

Responsibilities

  • Collaborate with scientists and engineers on diverse projects supporting mission-critical technical analysis for national space assets
  • Lead cross-functional teams and mentor junior engineers.
  • Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.
  • Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance.
  • Deliver high-quality HPC infrastructure design, and system configuration.
  • Develop and deploy automation solutions using tools such as Clush.
  • Manage infrastructure using Infrastructure-as-Code and GitOps practices
  • Implement, support, and optimize GPU computing.
  • Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency.
  • Develop cost-efficient HPC service offerings that align with mission and business objectives.
  • Harden Linux systems to meet stringent security requirements

Benefits

  • Comprehensive health care and wellness plans
  • Paid holidays, sick time, and vacation
  • Standard and alternate work schedules, including telework options
  • 401(k) Plan — Employees receive a total company-paid benefit of 8%, 10%, or 12% of eligible compensation based on years of service and matching contributions; employees are immediately eligible and vested in the plan upon hire
  • Flexible spending accounts
  • Variable pay program for exceptional contributions
  • Relocation assistance
  • Professional growth and development programs to help advance your career
  • Education assistance programs
  • An inclusive work environment built on teamwork, flexibility, and respect
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service