Member of Technical Staff - ClusterMAX

SemiAnalysis•San Francisco, CA
•Hybrid

About The Position

SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to state-of-the-art AI Models, CUDA kernels, and GPU cloud infrastructure. We are recognized as the leading authority on AI infrastructure, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies. We’re a global team of over 20 analysts & engineers, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually. Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies. ClusterMAX™ is the industry-standard GPU cloud rating system — 95% market coverage by volume, 84 providers rated, 209 tracked, and 140+ customer interviews behind each release. We recently published our cluster TCO and goodput framework in How Much Do GPU Clusters Really Cost? , showing that goodput expense alone swings 6–21% of total cluster TCO depending on fault-tolerance approach. We are now testing providers for ClusterMAX 3.0 with expanded benchmarks, security requirements, and analysis. As an MTS on ClusterMAX, you will build the benchmarks and run the evaluations that determine how the world's GPU clouds get rated.

Requirements

  • Hands-on experience operating GPU clusters: Slurm and/or Kubernetes, InfiniBand/RoCE fabrics, distributed storage
  • Strong Python and shell scripting; comfort building benchmark harnesses and CI pipelines
  • Understanding of distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery

Nice To Haves

  • Experience similar to our Technical Consultant profile is a plus: due diligence, TCO analysis, and client-facing technical communication
  • Security mindset: multi-tenant isolation, bare-metal vs virtualized trade-offs

Responsibilities

  • Leading development of the next generation of ClusterMAX™ benchmarks: storage IO and bandwidth, NCCL/RCCL collectives, fault-tolerance and goodput measurement, multi-tenant security and isolation testing
  • Deploying to and evaluating dozens of GPU clusters across hyperscalers and neoclouds (GB300/GB200 NVL72, B300/B200, H200, MI355X, TPUv7)
  • Extending our TCO and goodput methodology (see the ClusterMAX TCO & Goodput calculator ) into automated, reproducible tests
  • Working with executives and engineers at 200+ neoclouds, hyperscalers, and chip vendors
  • Authoring technical research analyzing benchmark results, reliability, security, and ease of use, with direct authorship recognition

Benefits

  • generous PTO
  • office stipends
  • competitive healthcare (medical, dental, vision)
  • support for conferences and ongoing learning
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service