Reinforcement Learning Infrastructure Engineer

ElorianPalo Alto, CA
$275,000 - $475,000Onsite

About The Position

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast. We're looking for an infrastructure engineer to design and build the core systems behind how we train our models with reinforcement learning (RL). You'll own the training infrastructure end to end, from rollout and reward pipelines to orchestration, reliability, and observability. The work spans both the algorithmic side of RL and the systems reality of running distributed training at scale, and you'll partner closely with our research team to keep RL training fast, stable, and dependable for the multimodal, visual reasoning models at the center of our work.

Requirements

  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)
  • Experience with actor-learner architectures and environment rollout orchestration at scale
  • Strong Python skills, plus PyTorch or JAX
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)
  • A track record of improving training throughput and GPU utilization at scale
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases

Nice To Haves

  • Experience with multimodal or agentic RL environments
  • Experience with RLHF or reward modeling pipelines
  • A self-directed builder who moves quickly and works across teams in an early-stage setting

Responsibilities

  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines
  • Improve GPU utilization and training throughput across the cluster

Benefits

  • health, dental, and vision benefits
  • unlimited PTO
  • paid parental leave
  • relocation support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service