About The Position

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

Requirements

  • 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting.
  • Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment).
  • Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops.
  • Team leadership with schedule accountability across multiple concurrent builds or sites.

Nice To Haves

  • GPU cluster validation experience (NCCL/RCCL benchmarking).
  • Automation skills (Python, Ansible) applied to deployment.
  • Optics/link-layer debugging depth.
  • Experience with acceptance testing as a commercial gate (revenue-linked).

Responsibilities

  • Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.
  • Lead deployment engineering across concurrent cluster builds, through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling vendors.
  • Drive deployment velocity engineering: cut bring-up time per cluster through tooling, pre-staging, and defect-source elimination, and set the targets the team is measured against.
  • Own defect feedback loops to Network Engineering (design), Layer One (cabling quality), and vendors (hardware/optics RMA patterns), holding those partners accountable to resolution.
  • Define spares, test equipment, and deployment tooling requirements per site, and standardize them across sites.
  • Own cluster validation end to end: bandwidth/latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates and raise the bar on each as the fleet scales.

Benefits

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service