GPU Systems Engineer

Tower Research CapitalNew York, NY
$200,000 - $300,000Hybrid

About The Position

Trading and research at Tower Research Capital run around the clock and across the globe, powered by infrastructure designed, built, and operated by this team. As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind this work at serious scale, managing hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The role is broad by design, involving shaping the architecture of new AI clusters, profiling training jobs, and writing automation to maintain the fleet with minimal human intervention.

Requirements

  • 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
  • Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
  • Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
  • Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
  • Clear communication. You will work daily with researchers, engineers, and vendors.

Nice To Haves

  • Experience with the rest of the NVIDIA stack, such as NCCL and NVLink.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Benefits

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service