Research HPC Systems Administrator

New Mexico State UniversityLas Cruces, NM
Onsite

About The Position

Join New Mexico State University as our High-Performance Computing (HPC) Systems Administrator and play a critical role in advancing groundbreaking research across diverse disciplines. In this dynamic position, you'll manage and optimize cutting-edge computing infrastructure, support researchers tackling complex computational challenges, and help shape the future of scientific discovery. If you're passionate about automation, large-scale computing environments, and empowering innovation through technology, this is your opportunity to make a lasting impact.

Requirements

  • Associate's Degree + 9 years of relevant experience or a Bachelor's degree + 7 years of relevant experience.
  • Relevant research computing experience, internships, professional training, and industry-recognized computing or information technology certifications may be substituted for portions of the required experience, as appropriate.

Nice To Haves

  • Master's degree or higher preferred.

Responsibilities

  • Administer, maintain, and optimize NMSU’s research high-performance computing (HPC) cluster, including compute nodes, login nodes, storage systems, networking interfaces, and supporting services.
  • Manage and tune the Slurm workload manager — job scheduling, partitions, QoS settings, node configurations, troubleshooting, and end-user support.
  • Oversee and maintain parallel file systems (PanFS), ensuring reliability, performance, and data integrity.
  • Monitor system performance and resource utilization, and identify and resolve performance bottlenecks.
  • Perform software installation and environment management (modules, conda, Spack) and apply system and security updates using modern automation tools (e.g., Ansible, Puppet, Terraform).
  • Develop documentation, training materials, and workshops to help researchers adopt the cluster effectively, and assist with building, deploying, and scaling containerized workloads (Singularity/Apptainer, Docker).
  • Ensure system security, compliance, backups, monitoring, and data-protection best practices, and assist with integrating the HPC environment into campus identity, networking, monitoring, and storage systems.
  • Contribute to long-term research-computing capacity planning and the procurement of new technology.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service