About The Position

This role is focused on designing and operating distributed training systems for large neural networks across GPU clusters. The engineer will optimize multi-node, multi-GPU execution, diagnose and resolve bottlenecks, and improve training stability and fault tolerance at scale. This position involves partnering with research and applied ML teams to productionize large-model training pipelines, ultimately building the foundation for large-scale AI.

Requirements

  • Deep hands‑on experience with distributed systems or ML systems
  • Experience running large‑scale workloads on GPU clusters
  • Production experience with PyTorch distributed training
  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
  • Low‑level understanding of GPU communication and networking
  • GPU orchestration: Slurm, Kubernetes, Ray, RunAI
  • Communication libraries: NCCL, RDMA, InfiniBand, NVLink
  • Training frameworks: PyTorch Distributed, Megatron‑LM, DeepSpeed
  • Memory optimisation: activation checkpointing, ZeRO offload techniques

Nice To Haves

  • Experience working with large language models or foundation models is a strong plus

Responsibilities

  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters
  • Optimize multi‑node, multi‑GPU execution to maximize throughput and utilization
  • Diagnose & resolve bottlenecks across compute, memory, and network
  • Improve training stability and fault tolerance at scale
  • Partner with research and applied ML teams to productionize large‑model training pipelines
  • Build and optimize GPU cluster orchestration using Slurm, Kubernetes, Ray, RunAI
  • Ensure efficient scheduling, isolation, and fairness across training workloads
  • Optimize and debug distributed communication using NCCL, RDMA, InfiniBand, NVLink
  • Minimize networking bottlenecks that dominate end‑to‑end training time
  • Scale large-model training using PyTorch Distributed, Megatron‑LM, DeepSpeed
  • Own multi‑node launch configurations, failure recovery, and performance tuning
  • Apply advanced memory optimization techniques: Activation checkpointing, ZeRO (Stage 1–3) and offload strategies
  • Balance compute, memory, and communication to push model size and batch scale
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service