ML Research Engineer, Training

Weave RoboticsSan Francisco, CA

About The Position

Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. We're one of the first companies with a deployed fleet generating real-world robot data at terabyte scale. The pipeline and training stack you build are what turns that data into capability. Model quality is set as much by training as by architecture: what data gets in, how it's sampled, whether the run is stable, whether a silent bug ate the gradient three days ago. You'll own that layer from raw fleet uploads to the batch that hits the GPU. When the stack is right, ideas become models in training in days, and deployed in weeks.

Requirements

  • ML System Expertise: Deep PyTorch or JAX experience, including multi-node distributed training (FSDP, DDP, or equivalent) on real workloads.
  • Performance engineering: Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance, including CUDA-level profiling tools (e.g. Nsight Systems) and NCCL tuning.
  • Training Run Judgement: You can read a loss curve, tell instability from a data bug and know when to kill a run.

Nice To Haves

  • Robot learning exposure: you’ve trained policies (VLAs, world models, RL) and can tell a data problem from a model problem.
  • Cluster and cloud infrastructure experience: Kubernetes, SLURM, GCP/AWS.
  • Large-scale post-training experience: SFT, reward modeling, RL fine-tuning.
  • On-robot inference optimization experience: TensorRT, quantization, distillation.
  • CUDA or Triton kernel work.
  • Experience with video-heavy datasets: transcoding, chunking, and the storage/compute tradeoffs of training on video at scale.

Responsibilities

  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
  • Large-scale data handling: Develop high-throughput data ingestion, transformation, and storage systems capable of processing terabytes of multimodal robot data, including video, proprioception, and sensor streams.
  • Optimize research productivity: Grow the codebase that makes experiments reproducible, scalable, and easy to launch, monitor, retry/recover, and debug.
  • Sampling and Curation: Drive sampling and curation decisions that show up in model behavior.
  • Make runs fast and honest: profile and fix throughput bottlenecks, chase down loss spikes and silent data bugs, keep results reproducible enough to trust: from data loading to GPU kernels.
  • From research to production: Turn research prototypes into infrastructure the whole team trains on.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service