HPC Infrastructure & Cluster Engineer

General Dynamics Information Technology•Springfield, VA
•$119,850 - $162,150•Onsite

About The Position

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract at GDIT. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Requirements

  • Active TS/SCI with the ability to obtain CI Poly.
  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
  • US Citizenship Required

Nice To Haves

  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Responsibilities

  • Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Benefits

  • 401K with company match
  • Comprehensive health and wellness packages
  • Internal mobility team dedicated to helping you own your career
  • Professional growth opportunities including paid education and certifications
  • Cutting-edge technology you can learn from
  • Rest and recharge with paid vacation and holidays
  • Variety of medical plan options, some with Health Savings Accounts
  • Dental plan options
  • Vision plan
  • 401(k) plan offering the ability to contribute both pre and post-tax dollars up to the IRS annual limits and receive a company match.
  • Full flex work weeks where possible
  • Variety of paid time off plans, including vacation, sick and personal time, holidays, paid parental, military, bereavement and jury duty leave.
  • Short and long-term disability benefits
  • Life, accidental death and dismemberment, personal accident, critical illness and business travel and accident insurance are provided or available.
Ā© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service