Research Platform Engineer

Tower Research Capital•New York, NY
•$200,000 - $300,000•Hybrid

About The Position

As part of Tower Research's Core Engineering team, you will develop a multi-tenant research compute platform capable of dynamically orchestrating large scale ML workloads across a hybrid compute infrastructure of GPUs and CPUs. Your primary mission is to work closely with Quant Researchers, Portfolio Managers, and Infrastructure teams, to build a unified, elastic, and multi-tenant research substrate.

Requirements

  • Deep experience with HPC job schedulers (e.g., Slurm), including gang scheduling, fair-share priority trees, topology-aware node allocation, and containerized HPC execution
  • Advanced understanding of Kubernetes architecture, CRDs, HPC focused operators, admission controllers and GPU-native schedulers
  • Familiarity with distributed computing frameworks (e.g., Ray) and deep learning frameworks (e.g., PyTorch, JAX) with multi-node scaling primitives
  • Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks
  • Experience designing vendor-agnostic infrastructure layers, cloud-bursting strategies, and compute execution spanning multiple on-premise datacenters and public cloud environments
  • Familiarity with high-performance storage solutions and high-speed network fabrics for large scale research workloads
  • Proven track record building cluster-wide telemetry and cost-attribution platforms for multi-tenant environments
  • A strong platform engineering mindset focused on reducing friction for researchers while maintaining rigorous operational efficiency, cost transparency, and system scalability

Responsibilities

  • Design intuitive platform abstractions, APIs, scheduler wrappers, and a durable job control plane so researchers can seamlessly launch simulations, distributed training runs, and complex research pipelines without incurring infrastructure overhead
  • Build a multi-tenant compute substrate combining HPC-grade batch scheduling (gang scheduling, topology awareness, fair share) with cloud-native flexibility, enforcing strict tenant isolation across trading desks
  • Establish infrastructure agnostic execution abstractions that enable compute workloads to run transparently across multiple on-premise datacenters or burst into external cloud providers
  • Evaluate, select, and integrate production-grade job graph orchestration engines to automate multi-stage feature, training, and backtesting pipelines
  • Implement automated failure detection, retry-on-fault mechanisms, and high-performance checkpointing to ensure long-running distributed jobs survive hardware degradation and faults without manual intervention
  • Profile and eliminate I/O bottlenecks, ensuring distributed ML workloads align with high-speed network fabrics and high-performance storage
  • Implement comprehensive telemetry to track compute utilization, queue pressure, and GPU/CPU cost attribution, giving Portfolio Managers and Senior Management clear visibility into resource consumption and ROI

Benefits

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service