ML Infra Engineer, Platform

Physical Intelligence•San Francisco, CA

About The Position

The Infrastructure team builds and operates the backbone of everything PI does: from training state-of-the-art VLA models, to orchestrating large-scale simulation, to reliably deploying intelligence across fleets of physical robots. The team works closely with researchers, robotics runtime, product, and platform engineers to ensure infrastructure scales from prototype to production-grade deployments.

Requirements

  • Extremely strong first-principles thinking.
  • Familiarity with agent infrastructure, observability, security and sandboxing
  • Deep experience with cloud platforms (GCP, AWS) and distributed systems: compute orchestration, networking, autoscaling, service meshes, load balancing.
  • Ability to reason about system bottlenecks, performance tuning, and cost optimizations across compute, networking, and storage.
  • Comfort with Kubernetes, cluster-level reliability, and service-oriented architectures.
  • Solid intuition around scalability, performance, and failure modes.
  • Experience with infrastructure-as-code (e.g. Terraform), containerization, and modern platform engineering practices.
  • Familiarity with logging, metrics, tracing, incident response, SLOs, and debugging complex distributed systems.
  • Strong cross-functional communication and ownership mindset.
  • Experience (4-6 years) working in fast-moving or early-stage environments where ambiguity is normal with demonstrated growth trajectory.
  • Ability to spin up quickly on unfamiliar and ambiguous domains, has a strong sense of ownership, and cares deeply about our mission.

Nice To Haves

  • Experience with large-scale ML training, evaluation, or simulation infrastructure.
  • Experience with secrets management systems (e.g., Doppler).
  • Background in observability, cost optimization, or internal platform tooling.
  • Exposure to robotics, simulation, or real-time systems.

Responsibilities

  • Own and scale AI-native infrastructure: Operate and evolve Kubernetes clusters and service deployment patterns, and help build a scalable microservice platform for internal systems such as evaluation services, operational tooling, and internal APIs with agent-use as the primary interface. This includes supporting safe rollouts, upgrades, and rollback strategies.
  • Drive security, observability and cost-aware infrastructure: Treat security, logging, metrics, tracing, and alerting as first-class platform primitives, and build systems that surface reliability and performance issues early. Also, help improve cost visibility and enable cost-aware decision-making at the infrastructure level.
  • Harden platform foundations: Own core infrastructure with multi-cloud considerations, designing authentication/authorization flows, networking architecture, quota and rate-limiting services, and cloud primitives that behave predictably. A major part of this work is reducing infra churn by standardizing patterns, abstractions, and interfaces.
  • Improve developer experience: Build clear, documented interfaces for using platform infrastructure, reducing the gap between “I need infra” and “I can run my workload.” This includes supporting consistent local vs. remote development workflows and improving self-serve infrastructure usage.
  • Collaborate and lead through ownership: Work closely with researchers and other engineers to understand requirements and constraints, translate fast-moving needs into reusable infrastructure, and own systems end-to-end, from design through operation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service