Staff Software Engineer - ML Infrastructure

WatneySan Francisco, CA

About The Position

Our Mission Expand human ambition in the physical world. Critical infrastructure is constrained by labor shortages, hazardous working conditions, and operational complexity. Watney builds and deploys autonomous robotic systems that increase the speed and capacity of buildout, starting with data centers. About the Role At Watney, ML Infrastructure engineers turn data collected from a live fleet of robots into better models. The fleet produces large volumes of video and telemetry data from real work in the field, and making that data trainable is one of the hardest systems problems at the company. As we continue to scale, these systems will require larger training runs with more data, expanded clusters, and optimal GPU utilization.

Requirements

  • Have built ML infrastructure that carried real production training runs
  • Have scaled distributed training systems
  • Strong experience with Python, PyTorch or TensorFlow
  • Have experience identifying and troubleshooting GPU performance bottlenecks in large-scale training environments

Responsibilities

  • Own training and inference infrastructure
  • Build the data pipelines that these training runs depend on
  • Make experiments fast to launch and reproduce
  • Contribute to our core training code
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service