This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day. The fleet has grown fast, with close to a thousand machines added recently, and this role exists so that infrastructure gets dedicated, high-standard operational care. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation. You will sit inside the HPC team, next to the engineers who build and run the platform. That proximity matters: your diagnostics feed their root-cause work, your runbooks capture what the team learns, and the recurring issues you surface become candidates for automation and permanent fixes. For someone who wants to grow into HPC engineering, this is a strong place to start.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level