Developer Relations Engineer

Andromeda ClusterSan Francisco, CA
Remote

About The Position

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster — but it filled almost instantly. Since then, we've been quietly building the systems, network, and orchestration layer that makes the world's AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

Requirements

  • Experience running distributed training or large-scale inference in production.
  • Experience debugging stalled multi-node runs, understanding the difference between fabric and dataloader problems.
  • Strong Python skills, with Go or Bash as a plus.
  • Experience building and maintaining production-grade tools and libraries, and comfort owning a repo others depend on.
  • Fluency across the modern stack: PyTorch, NCCL, containers, Kubernetes and/or Slurm, CUDA-adjacent tooling, and at least one serving framework (vLLM, SGLang, TensorRT-LLM, or equivalent).
  • Excellent technical writing skills: ability to explain complex topics clearly and concisely to a smart audience.
  • Public work that can be reviewed: blog posts, benchmarks, OSS contributions, owned documentation, or talks given.
  • Judgment about what to build: ability to identify artifacts that will unblock many teams.
  • Ability to ship in ambiguity: defining what good looks like and achieving it without a pre-established playbook or team.

Nice To Haves

  • Time embedded with external engineering teams: solutions architecture, professional services, deployed engineering, or technical account ownership at an infrastructure company.
  • Performance engineering depth: profiling distributed jobs, reasoning about MFU, interconnect behavior, and checkpoint I/O bottlenecks.

Responsibilities

  • Turn knowledge from engineering conversations into artifacts: benchmarks that hold up to scrutiny, reference architectures people deploy, documentation an engineer trusts, and a developer community where the answer is already there before someone has to ask us.
  • Write code for reference implementations of multi-node training on the platform, preflight scripts for first big jobs, and integrations for popular orchestration frameworks on heterogeneous supply.
  • Create technical content: Benchmarks, deep-dive posts, performance write-ups, postmortems worth publishing, and reference architectures for common training and serving stacks. Ensure every claim is reproducible and every number is measured.
  • Develop end-to-end developer documentation: Quickstarts, orchestration guides (Slurm, Kubernetes, direct SSH), storage and checkpointing patterns, and troubleshooting runbooks. Own the roadmap for documentation.
  • Foster the developer community: Manage public channels, GitHub presence, and office hours. Set the tone, answer hard questions, and make it a place experienced infra engineers want to be.
  • Build code that lowers the floor for new users: Example repos, container images, reference configs, and integrations with frameworks and schedulers.
  • Translate learnings from live workloads: Work with the solutions team on onboardings and incidents, collaborate with product engineers on fixes, and turn learnings into public artifacts.
  • Act as the developer's advocate internally: File hard bugs, argue for roadmap changes, and build missing pieces when it's the fastest path.
  • Enable partners and providers: Create integration guides, onboarding material, and technical explanations for building against the platform without a call.

Benefits

  • Competitive pay
  • Meaningful equity
  • Healthcare coverage for you and your dependents
  • Dental coverage for you and your dependents
  • Vision coverage for you and your dependents
  • 401(k)
  • Unlimited PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service