Member of Technical Staff, Infrastructure

Recruiting From ScratchSan Francisco, CA
Onsite

About The Position

Our client is building general-purpose autonomous robots designed to automate physical labor in real-world industrial environments. The company is developing the infrastructure required to train, deploy, and continuously improve AI models powering a fleet of physically deployed robots. The company has raised approximately $23M from leading investors and has assembled a highly technical team with backgrounds across leading AI, robotics, infrastructure, and technology companies. As a Member of Technical Staff, Infrastructure, you'll work on foundational systems spanning distributed computing, large-scale data pipelines, GPU infrastructure, model training, networking, deployment, and real-time robot operations. This is an opportunity for a high-slope systems engineer who wants broad ownership in an early-stage environment where infrastructure decisions have immediate impact and engineering standards are still being established.

Requirements

  • 2+ years of experience in infrastructure, distributed systems, or related engineering roles
  • Strong systems engineering background
  • Experience building production infrastructure or distributed systems
  • Experience working with large-scale data, compute, or infrastructure systems
  • Experience at a highly technical, tier-1 VC-backed startup, technology company, or research lab
  • Experience operating in environments with a high engineering talent bar
  • Experience owning technical projects end-to-end
  • Experience working in fast-moving startup or research environments
  • Strong ability to operate independently with minimal structure
  • Strong engineering judgment and problem-solving ability
  • Ability to move quickly between different technical problem areas
  • Strong communication skills and technical curiosity
  • Comfortable working on-site in San Francisco 6 days/week
  • Comfortable working long hours in a highly demanding startup environment
  • Strong distributed systems fundamentals
  • Strong systems engineering fundamentals
  • Experience building scalable infrastructure
  • Experience with large-scale data pipelines
  • Experience with cloud infrastructure and distributed compute
  • Experience with fault-tolerant distributed systems
  • Experience with networking and infrastructure systems
  • Experience with large-scale data processing
  • Strong Python programming skills
  • Experience working with production infrastructure
  • Strong understanding of system design and architecture
  • Ability to troubleshoot complex distributed systems
  • Ability to reason about performance, reliability, and scalability
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or related technical field preferred
  • Strong computer science and systems fundamentals
  • Equivalent practical engineering experience accepted
  • Exceptional ownership and execution ability
  • High technical slope and learning velocity
  • Strong engineering intuition
  • Highly hands-on engineering mindset
  • Comfortable operating with significant autonomy
  • Strong problem-solving ability
  • Strong technical communication skills
  • High energy and urgency
  • Comfortable working in highly demanding startup environments
  • Comfortable working across different technical domains
  • Strong builder mentality
  • Comfortable operating without established processes
  • Strong attention to engineering quality
  • Willing to tackle whichever technical bottleneck is most important
  • Comfortable making decisions with incomplete information
  • Strong collaboration skills
  • Comfortable working closely with highly technical engineers and researchers
  • Low-ego technical working style
  • Strong interest in AI, robotics, infrastructure, and distributed systems
  • Comfortable working 6 days/week onsite in San Francisco

Nice To Haves

  • Experience with GPU or high-performance computing infrastructure preferred
  • Experience with ML infrastructure or model training systems preferred
  • Experience with compute scheduling or orchestration preferred
  • Strong C/C++ programming skills or experience preferred
  • Experience with large-scale training infrastructure preferred
  • Experience with petabyte-scale data systems preferred
  • Experience with multi-node training orchestration preferred
  • Experience with real-time or low-latency systems preferred
  • Experience with robotics or autonomous systems preferred

Responsibilities

  • Build and own distributed systems infrastructure powering a fleet of deployed autonomous robots
  • Design infrastructure for large-scale model training and inference workloads
  • Build training orchestration systems for large-scale ML workloads
  • Design compute scheduling and resource allocation systems
  • Build fault-tolerant infrastructure for production AI and robotics workloads
  • Develop networking infrastructure supporting physically deployed robots
  • Design and maintain large-scale data pipelines for robot telemetry and video data
  • Build systems capable of ingesting and processing petabyte-scale datasets
  • Ensure training infrastructure can continuously ingest and process production robot data
  • Architect continuous-learning infrastructure connecting deployed robots to model training systems
  • Build reliable pipelines that move production trajectories from robots back into training
  • Optimize data infrastructure to keep GPU workloads continuously supplied with data
  • Build and improve GPU cluster orchestration and compute infrastructure
  • Work on low-latency infrastructure supporting robot operations across long distances
  • Design systems capable of meeting demanding latency and reliability requirements
  • Build internal infrastructure and developer tooling that improves engineering velocity
  • Identify infrastructure bottlenecks and own solutions end-to-end
  • Work across different areas of the stack depending on the highest-leverage technical problem
  • Develop core libraries and infrastructure components used across the engineering organization
  • Improve observability, reliability, and operational tooling across distributed systems
  • Establish engineering standards and best practices across infrastructure
  • Work closely with ML researchers, robotics engineers, and software engineers
  • Support infrastructure powering large-scale model training and experimentation
  • Build systems that can reliably operate in real-world production environments
  • Make architectural decisions in a rapidly evolving technical environment
  • Operate with significant autonomy and ownership
  • Move quickly from identifying a bottleneck to designing and shipping a solution
  • Help shape the company's long-term infrastructure and distributed systems architecture

Benefits

  • Competitive startup equity
  • H-1B and OPT transfer support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service