ML Infrastructure Engineer

Transparent Search Group•Redwood City, CA
•Onsite

About The Position

ML Infrastructure Engineer Company: Dyna Robotics Location: Redwood City, CA (in office 5 days per week) Compensation: $220,000 - $350,000 + competitive equity Employment Type: Full-time Visa Sponsorship: Visa transfers (OPT, H-1B transfer) About Dyna Robotics Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across environments with commercial-grade performance. Its affordable, intelligent robotic arms are already deployed at customer sites in hospitality and restaurants, automating repetitive, stationary tasks. Founded in 2024 by repeat founders who previously built and sold Kaper AI to Instacart, with a team from Google DeepMind, Meta and Cruise, Dyna has about 130 people and has raised $143.5M from investors including NVentures, Samsung NEXT, Salesforce Ventures, First Round Capital and CRV. The Role Dyna Robotics is hiring an ML Infrastructure Engineer to own training infrastructure end to end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You will be the connective tissue between researchers and compute, and your work directly speeds the path from model to deployed robot.

Requirements

  • 5-7+ years as an infrastructure engineer, including leading technical projects in HPC or ML infrastructure
  • Built and maintained ML or data infrastructure on a team with a high talent bar
  • Deep PyTorch experience and hands-on large-scale distributed training (FSDP, ZeRO, failure recovery)
  • GPU performance optimization and profiling (CUDA, NCCL, Triton)
  • Genuine interest in robotics and physical AI
  • Ability to work in the Redwood City office 5 days a week

Nice To Haves

  • Robotics experience at startups or enterprise teams
  • Early-stage or founding infrastructure hire
  • Multimodal systems (video, audio, multimedia models)
  • DeepSpeed or Accelerate; model serving optimization and monitoring

Responsibilities

  • Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimization (ZeRO, FSDP).
  • Build researcher-friendly tooling and job scheduling (Kubernetes, SLURM) with fast iteration, automated retries and failure recovery.
  • Design high-throughput pipelines that ingest terabytes of multimodal robot data (video, proprioception, 3D signals) so GPUs never starve.
  • Build low-latency inference pipelines for real-time robot control using quantization, distillation and compilation (TensorRT, Triton).
  • Profile GPU utilization, I/O bottlenecks and memory fragmentation to maximize fleet performance.

Benefits

  • competitive equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service