HPC Platform Engineer

Numerical Algorithms Group,
$100,000 - $130,000Remote

About The Position

Join a Team Where You'll Keep Learning If you're an experienced Platform Engineer with a strong Linux and High Performance Computing background looking for your next technical challenge, this could be the opportunity you've been looking for. At X-ISS, part of n² Group, we deliver managed services and project expertise for customers running complex, large-scale High-Performance Computing (HPC) environments across scientific research, artificial intelligence and enterprise workloads. You'll join a highly skilled, collaborative team where knowledge is shared, ideas are welcomed and everyone is encouraged to keep developing their expertise. You'll work alongside experienced engineers, solve challenging technical problems and deepen your HPC specialism day-to-day. This is a hands-on engineering role where no two days are quite the same. You'll work directly with customers, helping them maintain and improve business-critical HPC platforms while developing specialist skills in a highly specialized area of infrastructure. If you're passionate about Linux, enjoy solving complex technical challenges and are excited by the opportunity to build specialist HPC expertise, we'd love to hear from you.

Requirements

  • A bachelor’s degree, or equivalent qualification, in computer science.
  • A minimum of six years' professional experience implementing and supporting Linux server solutions.
  • Experience using service desk or ticketing systems to a high standard.
  • Exceptional troubleshooting skills across Linux servers and networking within large-scale environments, with the ability to work through interconnected hardware, network and software components across on-premise and cloud infrastructure to identify and resolve root causes.
  • Expertise in HPC platform engineering, with practical experience deploying, automating, and maintaining compute infrastructure in both on-premise and cloud environments, including containerization (Docker, Kubernetes) and Linux/Rocky 8 administration.
  • Hands-on experience with HPC job scheduling and workload management (e.g., Slurm, PBS Pro, or LSF), and automating bare-metal cluster provisioning (e.g., Ansible, xCAT, Warewulf, or Foreman/Cobbler) alongside equivalent automation in cloud environments.
  • Experience supporting and utilizing continuous integration and delivery (CI/CD) pipelines, enabling seamless, repeatable deployments across hybrid on-premise and cloud infrastructure. Familiarity with GitLab and DevOps pipelines is a plus.
  • Proficient programming skills in scripting and automation languages, such as bash or Python, and infrastructure-as-code tooling (e.g., Ansible or Terraform), for automating operational tasks, system configuration, and cluster provisioning at scale.
  • Familiarity with core HPC infrastructure components, including parallel file systems (e.g., Lustre, GPFS/Spectrum Scale, or BeeGFS), high-speed interconnects (InfiniBand/RDMA), and environment/module management (e.g., Lmod, Spack, or EasyBuild).
  • Strong interpersonal and communication skills to effectively engage and collaborate with both technical and non-technical team members, ensuring clarity and understanding across all collaborators.

Nice To Haves

  • Familiarity with databases, including deployment, optimization, and automation. MySQL experience preferred.
  • Linux and/or networking certifications.
  • Hands-on experience with cloud and on-premise infrastructure monitoring and observability tools (e.g., Prometheus/Grafana alongside cloud-native monitoring stacks), to ensure system observability and performance tracking.
  • Experience with GPU-accelerated computing environments, including NVIDIA drivers, CUDA, and GPU scheduling, is a plus.

Responsibilities

  • Collaborate closely with other platform engineers as well as users, tailoring hybrid on-premise and cloud HPC infrastructure to meet their unique compute and storage needs.
  • Manage, update, and optimize HPC resources across on-premise clusters and cloud environments, looking for opportunities for simplification and cost optimization while performing regular updates, firmware maintenance, and security patches.
  • Support infrastructure changes which will maintain and achieve set standards for platform uptime across bare-metal and cloud-hosted resources.
  • Support architecture and design decisions — including cluster topology, scheduler configuration, storage architecture, and cloud-bursting strategy — by providing technical expertise and working collaboratively with the wider team to identify the right solutions for our customers.
  • Develop and maintain reliable and scalable infrastructure, automating provisioning end-to-end across on-premise and cloud environments, to support the continuous and safe delivery of software solutions that meet the evolving needs of our users.
  • Ensure system observability and reliability by implementing monitoring tools spanning compute, network, and storage layers, and establishing clear escalation mechanisms for incident response, ensuring systems meet operational performance standards.

Benefits

  • A competitive salary
  • 401k matching
  • Health, Vision, Dental, Life and Supplemental Insurance
  • HSA
  • Tuition reimbursement.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service