HPC Systems Engineer

Numerical Algorithms Group,
Onsite

About The Position

Are you an experienced High-Performance Computing (HPC) Platform Engineer who enjoys solving complex technical challenges while working alongside skilled, collaborative colleagues? Do you have the expertise to design, build, operate, and optimize HPC platforms that support demanding scientific and engineering workloads? If so, we'd love to hear from you. Joining nAG means becoming part of a long-established organization with a reputation for technical excellence. For over 50 years, we've been helping organizations solve complex scientific and engineering challenges through world-class technical software, numerical expertise, and High-Performance Computing solutions. We value collaboration, innovation, and sharing knowledge, and as our HPC Services team continues to grow, you'll have the opportunity to work alongside experienced HPC specialists, helping design, optimize, and support a large-scale HPC environment. This position is based full-time at our client's Houston location, where you'll work as part of the nAG HPC Services team supporting a large-scale HPC environment. You'll have the opportunity to make a real impact on the evolution of a large-scale HPC environment supporting critical scientific and engineering workloads. You'll design, deploy, operate, and support high-performance computing platforms that power some of the energy industry's most computationally demanding scientific and engineering workloads. From deploying new infrastructure and optimizing existing platforms to diagnosing complex performance issues and evaluating new technologies, you'll play a key role in the ongoing evolution of the HPC environment. You'll take ownership of the day-to-day performance, reliability, and continuous improvement of a large-scale HPC environment supporting critical scientific and engineering workloads. Working as part of our HPC Services team, you'll collaborate closely with computational scientists, researchers, domain specialists, technology partners, vendors, and globally distributed technical teams to deliver a reliable, secure, and high-performing HPC environment. You'll play an important role in supporting critical workflows such as seismic data processing, reservoir visualization, and well planning by ensuring the platform continues to perform at its best.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Information Systems, or a related discipline, or equivalent practical experience.
  • Minimum of five years' hands-on experience administering Linux-based production environments (e.g., RHEL, CentOS).
  • Minimum of five years' experience deploying, administering, and supporting production HPC environments.
  • Experience with HPC technologies including one or more: Parallel or distributed file systems (e.g., Lustre, GPFS), High-speed interconnects (e.g., InfiniBand, Omni-Path), HPC workload schedulers (e.g., Slurm, PBS Pro)
  • Experience supporting production infrastructure, including networking, storage, compute, installation, configuration, maintenance, upgrades, and troubleshooting.
  • Solid understanding of data centre operations, including networking, cooling, and power.
  • Strong communication skills and the ability to work effectively with computational scientists and technical stakeholders.
  • Strong Linux systems administration skills.
  • Experience programming or scripting using Bash, Python, C, or C++.
  • Experience with HPC monitoring, troubleshooting, and performance optimization.
  • Experience with infrastructure automation and configuration management tools.
  • Experience with package management tools such as Conda, Spack, or RPM.
  • Experience supporting HPC applications using MPI.

Nice To Haves

  • Experience supporting multi-user HPC environments at scale.
  • Experience implementing infrastructure changes and security controls within enterprise or global environments.
  • Experience installing, compiling, and supporting vendor and open-source software.
  • Experience deploying or supporting infrastructure in public cloud environments.
  • Experience with GitLab CI/CD or similar automation tools.
  • Experience with container technologies supporting HPC workloads.

Responsibilities

  • Configuring, optimizing, and managing HPC clusters, storage systems, and networking components to support performance, reliability, and scalability.
  • Supporting the day-to-day operation, maintenance, and continuous improvement of HPC environments.
  • Diagnosing and resolving hardware, operating system, networking, storage, and application issues across the HPC stack.
  • Implementing appropriate security controls and maintaining platform integrity.
  • Collaborating with data scientists, researchers, computational scientists, and domain specialists to support and streamline technical workflows.
  • Monitoring system health and performance, investigating bottlenecks, and identifying opportunities for optimization.
  • Planning and performing software and operating system installations, upgrades, patches, and platform improvements.
  • Evaluating new hardware, software, and emerging technologies to improve platform capability and performance.
  • Working with technology partners and vendors to resolve complex infrastructure issues and evaluate emerging technologies and product roadmaps.
  • Ensuring HPC platforms continue to meet user, project, and organizational requirements.
  • Working closely with users to resolve complex technical issues and help them maximize the performance of scientific and engineering applications.
  • Supporting large-scale parallel file systems and the storage infrastructure underpinning HPC workloads.

Benefits

  • 401(k) plan with company match up to 5%
  • health, dental, life, short-term and long-term disability insurance
  • 10 vacation days
  • paid sick days
  • maternity and paternity leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service