Linux Systems Engineer 2

Pacific Northwest National LaboratoryRichland, WA
Hybrid

About The Position

We are seeking a Linux Systems Engineer to operate and expand the Computing and Data Operations (CDO) Linux infrastructure and high performance computing clusters that support EMSL's scientific instruments. This role covers hands on hardware work, Linux systems management, configuration automation, and Slurm based cluster operations across CDO's compute clusters (Tahoma and Boreal) and storage systems, including Ceph, VastData, BeeGFS, Lustre, and the Aurora/HPSS archive. The ideal candidate brings strong experience with on premises or hybrid Linux environments, modern automation tooling, and the ability to work collaboratively within a small, mission focused engineering team.

Requirements

  • BS/BA and 2 years of relevant experience -OR- MS/MA -OR- PhD
  • Strong Linux administration background including installation, patching, tuning, and troubleshooting.
  • Hands on automation experience with Ansible (YAML, roles, playbooks) or other system configuration tool.
  • Experience with virtualization technologies or cloud platforms (VirtualBox, Proxmox, AWS, Azure), or hybrid deployments.
  • Experience with Databases such as MySQL, MariaDB, Postgres, or MongoDB.
  • Scripting skills (bash, Python).
  • Comfort and physical ability to perform hands-on server hardware work — diagnostics, firmware updates, racking, and cabling — on-site in a data center environment.
  • Familiarity with GitLab CI or Github Actions and Git-centric workflows.
  • Basic knowledge of containerization (Docker/Kubernetes) used in DevOps and ML/HPC contexts.
  • Experience deploying or supporting HPC systems or large Linux installations.
  • Experience administering an HPC job scheduler (Slurm preferred).
  • Experience with bare-metal provisioning tools such as xCAT or Warewulf.
  • Experience using linux system packaging tools to deploy, remove, and package software.
  • Experience with large-scale storage systems such as Ceph, VastData, Lustre, BeeGFS, or HPSS.
  • Experience building dashboards, alerting, or automation on top of monitoring stacks such as Prometheus, Grafana, or ELK.
  • Experience collaborating with AI on scripting, software development, and/or system administration work.

Nice To Haves

  • Degree in Computer Science, Computer Information Systems, or a related field.

Responsibilities

  • Install, configure, maintain, and troubleshoot Linux servers (RHEL/Rocky and derivatives) across CDO's HPC clusters and EMSL support infrastructure, including bare-metal provisioning.
  • Administer HPC cluster operations on Tahoma and Boreal, including Slurm scheduler configuration, node provisioning with Warewulf, InfiniBand networking, and hardware diagnostics.
  • Develop and maintain automation using Ansible, including playbooks, roles, and configuration pipelines.
  • Support and understanding of large-scale storage systems (Ceph, VastData, BeeGFS, Lustre, and the Aurora/HPSS archive) underlying HPC compute and EMSL instrument data.
  • Work with DevOps tooling and workflows to ensure reliable software and infrastructure deployments (GitLab CI/CD, Git‑based workflows).
  • Participate in on‑call or rotational support for mission‑critical systems.
  • Monitor compute, storage, and network health using tools such as Prometheus, Grafana, Nagios, and ELK.
  • Assist with containerized workloads (Docker, Kubernetes) where applicable to HPC operational workflows.
  • Perform hands-on hardware work on-site — racking, cabling, component replacement, and diagnostics — as a regular part of data center operations.
  • Document processes, procedures, and checklists to ensure operational consistency across environments.
  • Collaborate closely with development, research, and infrastructure teams to improve cluster performance, reliability, and automation.

Benefits

  • medical insurance
  • dental insurance
  • vision insurance
  • robust telehealth care options
  • several mental health benefits
  • free wellness coaching
  • health savings account
  • flexible spending accounts
  • basic life insurance
  • disability insurance
  • employee assistance program
  • business travel insurance
  • tuition assistance
  • relocation
  • backup childcare
  • legal benefits
  • supplemental parental bonding leave
  • surrogacy and adoption assistance
  • fertility support
  • company-funded pension plan
  • 401 (k) savings plan with company match
  • 120 vacation hours per year
  • ten paid holidays per year
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service