Cluster Engineer

STN IncSan Francisco, CA

About The Position

We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.

Requirements

  • 7+ years designing or operating large-scale Linux infrastructure.
  • 5+ years supporting production GPU clusters for AI or HPC workloads.
  • Demonstrated experience building multi-node GPU training environments from the ground up.
  • Deep expertise with distributed PyTorch training.
  • Extensive experience troubleshooting and optimizing NCCL communications.
  • Strong understanding of distributed AI communication patterns, including: AllReduce, ReduceScatter, AllGather, Broadcast, Point-to-point communications.
  • Experience benchmarking distributed training using tools such as: nccl-tests, NVIDIA DCGM, Nsight Systems.
  • Strong understanding of GPU memory management, including: KV Cache, Activation checkpointing, Tensor Parallelism, Pipeline Parallelism, Data Parallelism.
  • Experience optimizing LLM inference throughput, including: Tokens/sec optimization, Batch sizing, Continuous batching, KV cache tuning, Memory bandwidth optimization.
  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
  • Expert-level Linux systems administration skills.
  • Experience with Slurm workload manager.
  • Experience using Pyxis and Enroot for containerized GPU workloads.
  • Strong scripting skills using Python and Bash.

Nice To Haves

  • MLPerf (preferred)
  • Triton (preferred)
  • TensorRT-LLM (preferred)
  • Experience deploying AI workloads on Kubernetes.
  • Experience with NVIDIA GPU Operator.
  • Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
  • Familiarity with MLPerf benchmarking.
  • Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
  • Experience automating infrastructure using Ansible, Terraform, or similar tools.
  • Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.

Responsibilities

  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
  • Build and support production AI infrastructure running hundreds to thousands of GPUs.
  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
  • Configure and tune distributed AI software stacks including: PyTorch, NCCL, CUDA, UCX, MPI, Slurm, Pyxis/Enroot.
  • Optimize GPU scheduling and resource allocation for both training and inference environments.
  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
  • Identify performance regressions and troubleshoot distributed training issues at scale.
  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
  • Work closely with ML engineers to improve training scalability and inference efficiency.
  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.
  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service