Member of the Technical Staff - Platform

Andromeda ClusterSan Francisco, CA
Hybrid

About The Position

Compute is the most sought-after resource in the world, yet it still trades like commercial real estate: year-long contracts, manual fulfillment, capacity sitting idle because nobody can move it. We are building the liquidity layer at Andromeda. Andromeda was founded by Nat Friedman and Daniel Gross to give startups the scaled AI infrastructure once reserved for hyperscalers. The first cluster filled almost instantly. The years since went into the platform that makes compute liquid: it deploys into foreign datacenters and turns the hardware it finds into clusters that leading AI labs train on. Today Andromeda operates compute for 80+ customers across 30+ capacity providers, with tens of thousands of GPUs under management and billions of GPU-hours supported, on everything from A100 to GB300.

Requirements

  • Impressive technical work you can go deep on, with impact in the world. That can take three years or twenty.
  • 2+ years of on-call experience for critical production services.
  • Deep Kubernetes experience.
  • Strong Linux fundamentals: kernel, cgroups, containers, networking, storage.
  • Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale.

Nice To Haves

  • Familiarity with fleet management and capacity planning.
  • Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware.

Responsibilities

  • Build and operate the control plane that runs our fleet.
  • Develop automated systems that take a cluster from bare machines to customer-ready.
  • Manage machine lifecycle between tenants: join, wipe, verify, rejoin.
  • Operate Kubernetes and Postgres across the fleet.
  • Contribute to our custom Kubernetes operators.
  • Scale clusters from tens of nodes to thousands.
  • Participate in on-call rotations.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service