HPC Linux System Administrator

MRI TechnologiesHouston, TX
Hybrid

About The Position

At NASA's Johnson Space Center, the Flight Sciences Laboratory (FSL) is the computing backbone behind nearly every major human spaceflight program - ISS, Orion, SLS, Commercial Crew, Lunar Gateway, and the Human Landing System: 700+ machines, 26,000 cores, 10+ petabytes of storage, 1,000+ users. MRI Technologies, supporting NASA under the JETS II contract, is hiring an HPC Linux System Administrator to help run and improve that cluster - a hands-on, on-premises HPC role administering the job scheduler and parallel filesystem, supporting containerized HPC workflows and CI/CD (including Jacamar), and working directly with the scientists and engineers behind NASA's human spaceflight mission. This is not a public-cloud, DevOps-only, or Kubernetes-only position. Your HPC experience doesn't have to say "NASA" on it. Another NASA center, a national lab, a university/research computing center, or another government agency all count - and your title doesn't have to have been "System Administrator." We want someone who'll bring fresh ideas from wherever they've been. Based in Houston, TX at NASA Johnson Space Center - fully remote considered for the right candidate with strong, demonstrated HPC experience.

Requirements

  • Minimum 5 years of Linux system administration in an on-premises, bare-metal, or cluster/research-computing environment.
  • Hands-on, production experience ADMINISTERING an HPC job scheduler (Slurm, PBS/Torque, or LSF).
  • Hands-on, production experience ADMINISTERING a high-speed parallel filesystem (Lustre or GPFS).
  • Experience using containers in an HPC context - packaging and running user environments (including older OS or older package versions) on a shared cluster rather than for traditional microservices.
  • Experience building and supporting CI/CD workflows, ideally tied to HPC clusters and run nodes.
  • System configuration management experience.
  • Experience with monitoring and alerting systems.
  • Demonstrated problem-solving, planning, and communication skills.
  • Ability to work effectively in a team environment.
  • Must be able to provide proof of U.S. Citizenship or U.S. Permanent Residency and complete a U.S. government background investigation.

Nice To Haves

  • Experience with RedHat-based Linux distributions.
  • Familiarity with InfiniBand high-speed networking.
  • Experience with provisioning tools (xCAT, Warewulf).
  • Experience with Ansible and/or Foreman for configuration management.
  • Familiarity with SPACK software package manager.
  • Experience with log consolidation, monitoring, and Git/GitLab (including CI/CD pipelines).
  • Familiarity with Jacamar (https://gitlab.com/ecp-ci/jacamar-ci) or comparable tooling for running CI/CD jobs against HPC clusters and run nodes.
  • Experience applying AI in a sysadmin role - integrating AI tooling into operational workflows, automation, and user-facing HPC jobs.
  • MPI workflow and administration experience (e.g., Open MPI, MPICH, Intel MPI).
  • Package management and environment modules experience (e.g., Lmod/Environment Modules, SPACK, EasyBuild).
  • Knowledge of NASA security mechanisms (security plans, POAMs, ATOs, Risk Assessments).

Responsibilities

  • Administering the job scheduler (Slurm, PBS/Torque, or LSF) including queue/partition configuration, accounting, and troubleshooting job failures.
  • Administering a high-speed parallel filesystem (Lustre or GPFS).
  • Supporting containerized HPC workflows and CI/CD (including Jacamar).
  • Working directly with scientists and engineers.
  • Building and supporting CI/CD workflows, ideally tied to HPC clusters and run nodes.
  • System configuration management.
  • Monitoring and alerting systems.

Benefits

  • medical
  • dental
  • vision
  • life and disability insurance
  • paid time off
  • 401(k)
  • 9/80 work schedule (every other Friday off, when applicable)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service