HPC Operations Engineer

Tower Research CapitalNew York, NY
$175,000 - $225,000Hybrid

About The Position

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day. The fleet has grown fast, with close to a thousand machines added recently, and this role exists so that infrastructure gets dedicated, high-standard operational care. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation. You will sit inside the HPC team, next to the engineers who build and run the platform. That proximity matters: your diagnostics feed their root-cause work, your runbooks capture what the team learns, and the recurring issues you surface become candidates for automation and permanent fixes. For someone who wants to grow into HPC engineering, this is a strong place to start.

Requirements

  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to escalate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.

Nice To Haves

  • Enough Bash or Python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.

Benefits

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service