HPC Linux System Administrator

MRI TechnologiesHouston, TX
Hybrid

About The Position

At NASA's Johnson Space Center, the Flight Sciences Laboratory (FSL) is the computing backbone behind nearly every major human spaceflight program - ISS, Orion, SLS, Commercial Crew, Lunar Gateway, and the Human Landing System: 700+ machines, 26,000 cores, 10+ petabytes of storage, 1,000+ users. MRI Technologies, supporting NASA under the JETS II contract, is hiring an HPC Linux System Administrator to help run and improve that cluster - a hands-on, on-premises HPC role administering the job scheduler and parallel filesystem, supporting containerized HPC workflows and CI/CD (including Jacamar), and working directly with the scientists and engineers behind NASA's human spaceflight mission. This is not a public-cloud, DevOps-only, or Kubernetes-only position. Your HPC experience doesn't have to say "NASA" on it. Another NASA center, a national lab, a university/research computing center, or another government agency all count - and your title doesn't have to have been "System Administrator." We want someone who'll bring fresh ideas from wherever they've been. Based in Houston, TX at NASA Johnson Space Center - fully remote considered for the right candidate with strong, demonstrated HPC experience.

Requirements

  • Hands-on, production experience administering an HPC job scheduler (Slurm, PBS/Torque, or LSF).
  • Hands-on, production experience administering a high-speed parallel filesystem (Lustre or GPFS).
  • Minimum 5 years of Linux system administration in an on-premises, bare-metal, or cluster/research-computing environment.
  • Bachelor's degree or equivalent certification in a related field, with a minimum of 5 years of experience.
  • Experience using containers in an HPC context.
  • Experience building and supporting CI/CD workflows, ideally tied to HPC clusters and run nodes.
  • System configuration management experience.
  • Experience with monitoring and alerting systems.
  • Demonstrated problem-solving, planning, and communication skills.
  • Ability to work effectively in a team environment.
  • Must be able to provide proof of U.S. Citizenship or U.S. Permanent Residency and complete a U.S. government background investigation.

Nice To Haves

  • Experience with RedHat-based Linux distributions.
  • Familiarity with InfiniBand high-speed networking.
  • Experience with provisioning tools (xCAT, Warewulf).
  • Experience with Ansible and/or Foreman for configuration management.
  • Familiarity with SPACK software package manager.
  • Experience with log consolidation, monitoring, and Git/GitLab (including CI/CD pipelines).
  • Familiarity with Jacamar or comparable tooling for running CI/CD jobs against HPC clusters and run nodes.
  • Experience applying AI in a sysadmin role.
  • MPI workflow and administration experience (e.g., Open MPI, MPICH, Intel MPI).
  • Package management and environment modules experience (e.g., Lmod/Environment Modules, SPACK, EasyBuild).
  • Knowledge of NASA security mechanisms (security plans, POAMs, ATOs, Risk Assessments).

Responsibilities

  • Administering the job scheduler (Slurm, PBS/Torque, or LSF) including queue/partition configuration, accounting, and troubleshooting job failures.
  • Administering a high-speed parallel filesystem (Lustre or GPFS).
  • Supporting containerized HPC workflows and CI/CD (including Jacamar).
  • Working directly with scientists and engineers.
  • Running and improving the HPC cluster.

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • Disability insurance
  • Paid time off
  • 401(k)
  • 9/80 work schedule (every other Friday off, when applicable)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service