Research Scientist, Scaling RL

Periodic Labs•Menlo Park, CA
•$250,000 - $350,000•Onsite

About The Position

We're training frontier models to develop deep scientific knowledge and reasoning for scientific tasks. You’ll study how RL scales with training compute, develop better algorithms, and take ideas from controlled experiments to our largest runs like Periodic Neon.

Requirements

  • Hands-on experience training LLMs with reinforcement learning
  • Strong attention to detail and rigorous approach to answer questions scientifically.
  • Coming up with small-scale RL setups that transfers to large-scale training runs.
  • Comfort working across a complex training stack to implement, debug, and test new research ideas.

Responsibilities

  • Design experiments to understand how RL performance scales with compute, model size, data, and reward quality, building on work such as ScaleRL
  • Develop better RL algorithms, spanning policy optimization, advantage estimation, exploration, and credit assignment for long-horizon RL tasks
  • Build adaptive sampling and curriculum methods that adjust task difficulty, problem selection, and the number of rollouts as models improve
  • Study bias and stability during RL training, including importance-sampling corrections and methods to tackle policy staleness and training–inference mismatch, as discussed here.
  • Improve compute efficiency across training and inference through experiments with hyperparameters, such as length penalties, rollout counts, batch sizes, and update schedules.

Benefits

  • Equity
  • Visa sponsorship
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service