HPC Operations Engineer

Evergreen Statistical TradingBellevue, WA
$200,000Onsite

About The Position

As an HPC Operations Engineer at Evergreen, you will own the day-to-day operation of our research clusters across multiple datacenters, along with the scheduled workloads that our research and trading operations depend on. You will also have the unique opportunity to work on a technology stack that is unencumbered by legacy infrastructure, and to take on substantial projects across both software and hardware as our HPC footprint continues to expand.

Requirements

  • Strong hands-on Linux systems administration experience, particularly with Red Hat-family distributions such as RHEL, Rocky Linux, AlmaLinux, etc.
  • Strong proficiency with Python for automation and tooling
  • Experience with version control tools, ideally Git, as well as versioned configuration management
  • Experience with monitoring and alerting systems, ideally Prometheus and Grafana
  • Experience with batch schedulers, ideally Slurm, including queue configuration, resource management and troubleshooting
  • Experience with cluster provisioning tools, such as Warewulf, xCAT, etc.
  • Experience with ZFS and parallel filesystems, such as GPFS, Lustre, etc.
  • Experience with InfiniBand networking
  • Ability to own critical processes, design them not to break, and resolve critical issues outside of business hours when they do
  • Ability to direct hardware work through datacenter remote hands, and be hands-on when the situation calls for it
  • Ability to diagnose complex problems in production systems under time pressure
  • Growth-oriented and collaborative mindset; enjoys working within a team
  • Passion for high-performance infrastructure that runs correctly and with minimal drama

Responsibilities

  • Own the day-to-day operation of research clusters across multiple datacenters.
  • Manage scheduled workloads that research and trading operations depend on.
  • Work on a technology stack unencumbered by legacy infrastructure.
  • Take on substantial projects across both software and hardware as the HPC footprint expands.
  • Direct hardware work through datacenter remote hands and be hands-on when necessary.
  • Diagnose complex problems in production systems under time pressure.

Benefits

  • Company-paid medical and/or other benefits
  • Signing and performance bonuses
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service