Head of Robot Learning

Anvil RoboticsTaipei, CA
Onsite

About The Position

This role is Anvil's first robot learning hire, responsible for implementing and deploying cutting-edge robot learning papers onto Anvil's hardware. The position involves taking research papers, getting them running on hardware within weeks, and producing polished demos and guides. Anvil is developing accessible robotics hardware and software, having already shipped over 200 robots and generated significant revenue. This role offers unique leverage over data collection, with the ability to influence the design of data collection systems and access to real-world industrial data through Anvil's manufacturing line and partner facilities. The successful candidate will be the sole member of the robot learning team initially, responsible for establishing the function and driving key product initiatives, such as closing the loop on training policies purely from handheld data collector output.

Requirements

  • Personally trained and deployed imitation-learning policies (ACT, Diffusion Policy, VLA fine-tunes) on real robot arms.
  • Fluent in the layer under the model: action chunking and temporal ensembling, inference latency versus control-loop frequency, camera synchronization and timestamp alignment, joint-space versus cartesian command interfaces.
  • Ability to replicate papers in weeks.
  • Strong data instincts to identify issues in teleoperation demonstrations.
  • Proven ability to finish projects with a guide, video, and reproducible repo.
  • Honest reporting of results, including n, eval protocol, and failure modes.
  • Comfortable being the only ML person in a fast, lean, founder-led company.
  • Ability to ship on a weeks-not-quarters cadence.
  • Master's or Bachelor's in CS, robotics, or a related field.
  • 2-4 years of fast-growing experience, or 1-2 years on a steep curve, as a research engineer working closely with a strong robot learning lead.
  • Experience making lab or team work run on hardware and closing hard problems personally.

Nice To Haves

  • PhD is explicitly not required or expected.
  • Publication record is not the bar.
  • Portfolio of policies personally trained running on real hardware, with repos and videos.
  • Experience as the person who made the lab's or team's work actually run on hardware.
  • Experience closing a growing share of hard problems personally.
  • Experience in the role of the person who never owned the direction because someone above them did.
  • Experience with LeRobot or similar open-source contributions.
  • Experience with DAgger / interactive imitation learning.
  • Experience with public demos that got real reach.
  • Experience with RL fine-tuning on real hardware.

Responsibilities

  • Train and deploy imitation-learning policies on real robot arms.
  • Understand and address the challenges in the layer beneath the model, including action chunking, temporal ensembling, inference latency, control-loop frequency, camera synchronization, timestamp alignment, and joint-space versus Cartesian command interfaces.
  • Replicate research papers and their associated codebases within weeks.
  • Identify and address issues in teleoperation demonstrations, such as inconsistent grasps, occlusions, and timing skew.
  • Ensure all shipped projects are polished, including a written guide, a video demonstration, and a reproducible repository.
  • Report results honestly, including sample size (n), evaluation protocols, and failure modes, especially for public releases.
  • Operate effectively as the sole ML person in a fast-paced, lean, founder-led company, setting personal agenda and shipping on a weeks-not-quarters cadence.
  • Develop and own training pipelines, from data ingestion and dataset formatting to training jobs and evaluation harnesses.
  • Replicate high-leverage public work (e.g., folding-class manipulation, VLA fine-tunes, diffusion policies) and publish it as a public demo and reproducible guide.
  • Validate the UMI pipeline, proving or fixing the path from the handheld data collector to a working policy on an OpenARM, and prescribing hardware revisions based on findings.
  • Build the data flywheel through DAgger/human-in-the-loop correction workflows and scale data collection beyond oneself by designing protocols for dedicated operators.
  • Maintain a high polish bar for all shipped projects, ensuring they include a video, a guide, and a repository with a README.

Benefits

  • Health and Wellness
  • Compensation and Support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service