Linux Cluster Hardware Administrator

The University of Texas at AustinAustin, TX
Onsite

About The Position

The Texas Advanced Computing Center (TACC) at The University of Texas at Austin is a leading supercomputing center supporting thousands of researchers and students. TACC staff assist researchers and educators in effectively utilizing advanced computing, visualization, and storage technologies, while also conducting research and development to enhance these technologies. Additionally, TACC staff educate and train future researchers. The Large Scale Systems group is seeking a hardware-focused Systems Administrator with extensive deployment experience in OmniPath 400Gbp/s and XDR Infiniband fabrics within a large-scale academic high-performance computing environment. Demonstrable experience with hardware upgrades from OmniPath 100Gbp/s to OmniPath 400Gbp/s is required. This role involves collaboration with various vendors and on-campus teams in a high-performance computing setting, spanning two physical locations with Petascale HPC/storage systems and other high-speed networking infrastructure. Candidates are required to submit a resume, letter of interest, and three references.

Requirements

  • 5+ years of hands-on experience with troubleshooting and repair of telecommunications equipment and computer hardware.
  • Experience with OmniPath 400Gbp/s in an HPC environment.
  • Experience with XDR Infiniband in an HPC environment.
  • Experience with Linux systems administration tools.
  • Experience with server hardware and physical repairs of high-density HPC systems.
  • Experience with physical deployments including cabling, power distribution, and familiarity with advanced datacenter cooling technologies (DLC and immersion cooling).
  • Ability and desire to learn new concepts and software tools.
  • Excellent verbal/written communication skills.
  • Must be eligible to work in the US on a full-time basis for any employer without sponsorship.

Nice To Haves

  • Experience in physical installation of large-scale clusters.
  • Experience with clusters that have multiple interconnect fabrics: multiple Infiniband technologies/OmniPath 100Gbp/s & OmniPath 400Gbp/s.
  • Working knowledge of systems provisioning techniques and tools.
  • Basic knowledge with one or more scripting languages (Python, Perl, Bash).
  • Working knowledge of monitoring software such as Nagios or Zabbix.

Responsibilities

  • Diagnose and repair Linux cluster hardware components, including standard datacenter server components, air-cooled components, Direct-Liquid-Cooling (DLC) infrastructure components (cold-plates, coolant distribution units - CDUs), Coolant-Oil-Immersion-Cooling (COIC) subsystem components (CDUs), and Petabyte-scale storage hardware components.
  • Maintain an up-to-date catalog of all work performed.
  • Deploy new systems and subsystems, including the deployment and management of high-speed networking fabrics.
  • Optimize rack positioning and internal layout to ensure power and thermal efficiency.

Benefits

  • 100% employer-paid basic medical coverage
  • Retirement contributions
  • Paid vacation and sick time
  • Paid holidays
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service