ML Infrastructure Engineer

ZiplineSouth San Francisco, CA
$160,000 - $250,000

About The Position

As an ML Training & Inference Infrastructure Engineer on the Data Platform team, you will be building and scaling the systems powering our data flywheel. This person will work at the intersection of autonomy and the infrastructure, owning systems that make ML development faster, reproducible, observable, and safe. This role is for a strong software engineer who enjoys the full ML development cycle: data ingestion, processing pipelines, dataset management, distributed training, continuous model integration, evaluation, and deployment. The ideal candidate has strong production engineering habits and is excited to build infrastructure that helps real autonomous systems improve over time.

Requirements

  • 3+ years of professional software engineering experience, ideally including ML infrastructure, data infrastructure, robotics, autonomy, aerospace, medical devices, or another safety-critical hardware/product environment.
  • Strong software engineering practices in Python in a production setting; comfort designing APIs, services, schemas, jobs, and operational workflows.
  • Experience building reproducible data pipelines and machine-learning pipelines.
  • Experience monitoring data statistics, system performance metrics, pipeline failures, and model/evaluation signals.
  • Working knowledge of ML concepts such as datasets, training, evaluation, optimization, statistics, and modern deep learning workflows.
  • Generalist mindset and willingness to work across cloud services, data platforms, developer tooling, and embedded/robotics-adjacent constraints.
  • Experience with PyTorch or similar ML frameworks.
  • Strong ownership, clear communication, and interest in building secure systems for mission-critical workflows.
  • Experience with Kubernetes or other container orchestration systems for production workloads.
  • Experience with cloud and on-premise production infrastructure, preferably AWS, and infrastructure-as-code tools such as Terraform or CloudFormation.

Nice To Haves

  • Experience deploying or evaluating ML systems on real robots, autonomous vehicles, drones, or other hardware products.
  • Experience with large-scale training systems, feature stores, data/versioned artifact stores, model registries, or experiment tracking.
  • Experience with annotation systems, dataset inspection tooling, or active-learning workflows.

Responsibilities

  • Build and operate software infrastructure that enables learning algorithms to leverage Zipline’s large-scale (quickly growing!) fleet data.
  • Design scalable, maintainable data and ML infrastructure for autonomy teams, including dataset creation, validation, training, evaluation, and deployment.
  • Own and improve data pipelines that feed into the ML development loop.
  • Identify and mitigate bottlenecks in the ML development cycle, especially around orchestration, performance, and reproducibility to increase the rate at which we can improve and scale the delivery experience.

Benefits

  • medical, dental and vision insurance
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service