Software Engineer, ML Infrastructure Platform

NuroMountain View, CA
$160,360 - $240,540

About The Position

Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 1+ years of relevant work experience.
  • Willingness to deep-dive into implementation and to raise the technical and operational standards of the broader engineering organization.
  • A demonstrated ownership mindset: you drive systems to operational maturity e.g. through monitoring, alerting, runbooks.
  • Strong proficiency in Python (and comfort with C++, Go or a similar systems language).
  • Hands-on experience running production infrastructure on Kubernetes.
  • Solid distributed-systems fundamentals and the ability to reason about performance, failure modes, and reliability across a complex system.

Nice To Haves

  • Strong working knowledge of GCP.
  • Experience with building large scale data generation pipelines.
  • Experience with Kubernetes-native orchestration for ML workloads.
  • Depth in GPU / distributed training internals, including NCCL and collective communication.
  • Familiarity with GPU and training observability tooling and using it to diagnose real bottlenecks.
  • A track record of driving down infrastructure cost while improving reliability.

Responsibilities

  • Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration.
  • Design and operate large-scale data pipelines - batch and streaming ingestion, storage layout, and high-throughput data generation and storage.
  • Design and develop agentic-first ML workflows - data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
  • Own reliability for critical training and release pipelines: instrument them, define meaningful alerting, and build the on-call and incident-response practices that let the team catch regressions.

Benefits

  • annual performance bonus
  • equity
  • competitive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service