Member of Technical Staff — Training

RadixArkPalo Alto, CA
Remote

About The Position

RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, accuracy and reliability across 10k, or 100k+ of GPUs. This role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact how next-generation AI models are trained and scaled. This is a deeply technical, high-impact role for engineers who enjoy solving hard systems problems at extreme scale.

Requirements

  • 3+ years of experience in ML systems, or large-scale training infrastructure
  • Experience building or operating large-scale agentic post-training systems.
  • Experience working on training / inference correctness or other precision-related problem
  • Experience debugging performance and stability issues in large post-training jobs
  • Experience improving training or inference efficiency.

Nice To Haves

  • Experience training 100+ billion-parameter models
  • Experience with train / inference optimization for large-scale RL or other production workload.
  • Familiarity with training stacks (e.g. Megatron-LM, FSDP, torchtitan, etc.) and inference stack (e.g. SGLang, vLLM, etc.)
  • Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL, AReaL, etc.)
  • Experience with RDMA, InfiniBand, NVLink, NCCL/RCCL, or high-speed GPU interconnects
  • Contributions to ML systems open-source projects
  • Experience with checkpointing, fault recovery, and elastic training.
  • Experience building infrastructure for agentic post-training, such as async rollout pipelines, sandbox, or harness system.

Responsibilities

  • Contribute to open-source large-scale post-training infrastructure Miles, and inference system SGLang.
  • Optimize throughput, scalability, and hardware efficiency
  • Improve reliability and fault tolerance for long-running training jobs
  • Develop training frameworks and infrastructure tooling
  • Collaborate with model researchers to support frontier experiments
  • Debug and resolve cross-layer performance bottlenecks
  • Build observability systems for training performance and reliability
  • Drive capacity planning and cluster utilization strategies
  • Contribute to long-term training infrastructure architecture

Benefits

  • competitive compensation with meaningful equity
  • comprehensive benefits
  • flexible work arrangements
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service