About The Position

An Emory College Research Support (ECRS) Senior Operating System Analyst/Administrator supports and administers Linux-based high-performance computing (HPC) systems with a primary focus on educational and instructional computing. Working with faculty, instructors, students, and other technical staff, this position helps provide reliable, accessible, and effective computing environments for courses that use Linux, HPC, GPU, and other computational resources. The position is responsible for the administration, maintenance, monitoring, and troubleshooting of Linux servers, HPC clusters, and related infrastructure. A significant part of the role involves administering technologies that provide students and instructors access to computational resources, including Open OnDemand and Slurm-based computing environments. This is a role intended for an administrator who can independently manage Linux systems and troubleshoot moderately complex infrastructure problems. The administrator is expected to possess a working knowledge of HPC concepts while continuing to develop expertise in cluster computing, storage, networking, automation, GPU computing, and other advanced technologies. The position may also support research computing infrastructure and collaborate with researchers whose computational requirements overlap with the technologies and services provided for educational computing.

Requirements

  • Five years of operating systems analysis/administration experience OR a bachelor's degree and three years of operating systems analysis/administration experience.
  • Professional experience administering Linux servers or comparable Linux infrastructure.
  • Demonstrated proficiency with the Linux command line and core Linux administration concepts, including filesystems, permissions, processes, services, package management, system logging, and system configuration.
  • Experience installing, configuring, maintaining, and troubleshooting Linux operating systems.
  • Working knowledge of TCP/IP networking, DNS, SSH, and network troubleshooting.
  • Experience administering multiuser Linux systems.
  • Experience with shell scripting and/or a general-purpose scripting language such as Bash or Python.
  • Experience diagnosing and resolving system administration problems independently.
  • Understanding of fundamental security practices for Linux systems, including patching, access control, privilege management, and vulnerability remediation.
  • Ability to document technical configurations, procedures, and troubleshooting information clearly.
  • Ability to communicate effectively with users possessing a wide range of technical experience.
  • Ability to manage multiple technical responsibilities and determine when issues require escalation or collaboration.
  • Demonstrated ability to learn and independently apply unfamiliar technologies.

Nice To Haves

  • A bachelor’s degree in computer science, information technology, engineering, a scientific discipline, or a related field is desirable, although equivalent of education and experience will also be considered.
  • Experience with Open OnDemand administration and configuration
  • Experience with Slurm workload management
  • Experience administering Linux-based HPC or cluster computing environments
  • Experience supporting computing resources used for classroom instruction
  • Experience with configuration-management and automation tools such as Ansible
  • Experience with Git or other version-control systems
  • Experience with LDAP or other centralized authentication and identity-management systems
  • Experience with GPU computing and scheduling GPU resources
  • Experience with NFS or other network filesystems
  • Experience with system and infrastructure monitoring
  • Experience with virtualization or container technologies
  • Experience with scientific or computational software environments
  • Experience with BeeGFS and other parallel filesystems
  • Experience with ZFS and advanced storage technologies
  • Experience with InfiniBand and other high-speed networking
  • Experience with advanced GPU resource management
  • Experience with HPC cluster architecture and performance
  • Experience with Splunk administration
  • Experience with configuration management and infrastructure automation
  • Experience with containers and application environments
  • Experience with scientific and research software stacks
  • Experience with hardware installation, diagnosis, and lifecycle management
  • Experience with system, application, and infrastructure monitoring
  • Experience with research computing environments

Responsibilities

  • Plans and implements one or more muli-platform operating systems, utilities, and related software to meet organizational needs.
  • May be responsible for applications on dedicated servers.
  • Ensures the availability, integrity and reliability of assigned systems.
  • Independently administer Linux systems used for instructional and educational computing.
  • Administer and maintain HPC clusters and other Linux-based computing environments used by students and faculty.
  • Administer, configure, maintain, and troubleshoot Open OnDemand environments providing web-based access to computational resources.
  • Administer and troubleshoot Slurm workload-management environments, including nodes, partitions, resource allocation, job scheduling, and common user issues.
  • Collaborate with faculty and instructors to prepare computing environments for courses, including determining resource requirements, configuring software, provisioning access, and testing environments before instructional use.
  • Provide advanced technical assistance to faculty, instructors, teaching assistants, and students using Linux and HPC resources.
  • Manage Linux user accounts, groups, permissions, SSH access, authentication, and integration with centralized identity-management systems.
  • Install, configure, patch, upgrade, and troubleshoot Linux operating systems and applications.
  • Install, configure, and maintain software required for instructional and computational workloads.
  • Support GPU-enabled computing environments and assist users with accessing GPU resources through the workload scheduler.
  • Perform operating system patching, security updates, vulnerability remediation, and other routine security administration.
  • Monitor system availability, utilization, capacity, and performance and proactively identify potential problems.
  • Diagnose and resolve hardware, operating system, storage, networking, authentication, scheduling, and application problems.
  • Develop and maintain system administration documentation, operational procedures, configuration standards, and user-facing documentation.
  • Use scripting and automation to improve system administration, deployment, configuration, and maintenance processes.
  • Participate in infrastructure upgrades, migrations, hardware deployments, and other computing projects.
  • Assist with evaluating new technologies and recommend improvements to instructional computing environments.
  • Provide technical guidance and assistance to junior administrators and student employees when appropriate.
  • Participate in the support of research computing systems when needed and apply experience gained across educational and research environments.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service