Senior Software Engineer, ML Infrastructure

ApptronikAustin, TX
Onsite

About The Position

Apptronik is building Apollo, a general-purpose humanoid robot, and the physical AI that drives it. Scale is the name of the game: every robot and teleoperator we field produces synchronized video, proprioceptive, tactile, and force-torque streams, and the fleet's output grows with every deployment. Turning that volume of data into shipped autonomy — routinely, at multi-terabyte scale — is what this role is about. We are looking for a Senior Software Engineer, ML Infrastructure to build that platform: the self-serve services and pipelines that carry data from collection through curation, training, and evaluation to a qualified model running on real hardware. Much of it is being created ground-up for the long term — humanoid robotics has few off-the-shelf answers — so the team builds first-party platform services alongside the open-source and commercial tooling we adopt where it genuinely fits. This is a hands-on role on a small team whose platform is depended on daily by researchers and engineers across MLOps, Autonomy, Data Platform, and TeleOp.

Requirements

  • A builder at scale: a track record of designing and shipping production systems and services that other teams depend on daily.
  • Deep hands-on experience with large-scale data pipelines for ML: multi-terabyte transformation and dataset assembly of multimodal sensor data — video and image streams, time-synchronized robot telemetry, the kind of data that trains vision-language-action and computer-vision models — with columnar and time-series formats (Parquet, Arrow), dataset versioning and lineage (lakeFS, DVC, Iceberg, or equivalent), and object storage (S3, MinIO).
  • Experience with ML annotation and labeling at scale: automatic annotation of data combined with human-in-the-loop workflows — the tooling, quality control, and throughput management.
  • Experience building large-scale evaluation or simulation harnesses: many parallel jobs on GPU infrastructure, aggregated into decision-grade results.
  • Strong Python and general software engineering ability (testing, API design, code review), plus cloud infrastructure, Kubernetes, Docker, and modern CI/CD.

Nice To Haves

  • Robotics data formats and fleet-scale telemetry (MCAP, ROS, LeRobot, or equivalent).
  • Simulation-in-the-loop evaluation with Isaac Sim, IsaacLab, MuJoCo, or equivalent.
  • Reinforcement or imitation learning infrastructure for embodied agents (rollout workers, sim-eval harnesses).
  • Deploying ML models to edge targets (ONNX Runtime, TensorRT, robot fleets).

Responsibilities

  • Build the ML platform — the APIs, workers, and control planes that let researchers and robot teams move data and models through the system in a self-serve manner, with the testing and observability that being a dependency implies.
  • Turn raw robot and simulation data into training-ready datasets — selection and filtering of manipulation episodes with synchronized sensor streams; annotation workflows that combine automatic labeling with human-in-the-loop review at throughput; and dataset versioning and lineage strong enough that any model traces back to the exact data that produced it.
  • Make multi-terabyte dataset operations routine — transformation and assembly, coverage and quality statistics that tell us a training set is good before we spend a cluster-week on it, and read paths that keep GPUs fed.
  • Build the rollout harnesses that evaluate policies in simulation on our GPU cluster; the benchmarks and metrics captured consistently across simulation, real-robot, and teleoperation sources; and the qualification gates a model must pass before it reaches Apollo — automatic, not manual review.
  • Build the model store — versioning, metadata, attached evaluation results, lineage — and the promotion path from trained to qualified to deployed on robot, including packaging (ONNX, TensorRT) in partnership with Autonomy.
  • Provide the tooling researchers use daily — experiment tracking, training job submission, sweeps, and reproducible container environments. Reduce time from idea to running training job; win adoption by being the fastest path, not by mandate.
  • Partner with Autonomy, Data Platform, and TeleOp on dataset and model lifecycle contracts.
  • Contribute to the technical direction of these layers.
  • Mentor the engineers around you through code and design review.

Benefits

  • Equal employment opportunities to all employees and applicants for employment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service