Senior/Staff Software Engineer, Infrastructure (ML)

Nimble RoboticsSan Francisco, CA

About The Position

We’re looking for a Software Engineer to join our ML Infrastructure team. In this role, you’ll help build the training and inference systems that power our general-purpose warehouse robots. You’ll own training infrastructure end to end: keeping GPUs highly utilized, making runs reproducible, and ensuring every researcher can launch the next experiment with a single command. You’ll work closely with ML and Robotics teams to design, build, and scale the systems that turn our GPU clusters into a reliable, high-throughput platform for model development.

Requirements

  • Bachelor’s, Master’s, or PhD in Computer Science or a related field, or equivalent practical experience.
  • 4+ years of industry experience in infrastructure, distributed systems, ML systems, robotics, or a related area.
  • Experience with programming languages such as Rust, Go, Python, or C++.
  • Experience with ML frameworks such as PyTorch or JAX.
  • Strong understanding of distributed systems, systems programming fundamentals, memory management, and performance optimization.
  • Experience with Kubernetes orchestration, resource scheduling for large distributed jobs, and containerized deployment pipelines.
  • Ability to debug and optimize bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations.
  • Ability to reason from first principles and optimize systems for both memory-bound and compute-bound workloads.
  • Strong cross-functional communication skills, ownership, and a growth mindset.

Nice To Haves

  • Hands-on experience with distributed training frameworks and techniques such as PyTorch DDP/FSDP, DeepSpeed, Megatron, or NCCL.
  • Hands-on experience with GPU kernel development.
  • Experience with data engineering technologies such as Parquet, Arrow, or similar systems.

Responsibilities

  • Design, develop, and maintain ML training infrastructure that enables the AI team to run training jobs efficiently, manage and iterate experiments quickly.
  • Build low-latency inference pipelines for production robotics workloads.
  • Develop, tune, and optimize low-level CUDA kernels.
  • Design training-platform systems for scalable model training, including high-throughput data ingestion, dataset sharding and sampling for distributed training.
  • Participate in and lead design reviews with peers and stakeholders to evaluate technical tradeoffs and select appropriate technologies.
  • Review code and provide feedback to uphold best practices around style, correctness, testability, performance, and maintainability.
  • Contribute to documentation and educational materials, adapting content as systems and workflows evolve.
  • Mentor junior engineers and help raise the technical bar across the team.

Benefits

  • Unlimited Flexible Time Off
  • Health Insurance (medical, dental, and vision)
  • Paid Parental Leave
  • Commuter Benefits
  • Referral Bonus
  • 401k
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service