Research Engineer, Infrastructure, RL Systems

Thinking Machines LabSan Francisco, CA
$350,000 - $475,000Onsite

About The Position

Thinking Machines is seeking an infrastructure research engineer to design and build the core systems for scalable, efficient training of large models using reinforcement learning. This role bridges research and large-scale systems engineering, requiring expertise in RL algorithms and distributed training/inference. The engineer will optimize pipelines, enhance reliability and observability, and collaborate with researchers and infra teams to make reinforcement learning production-ready. This is an evergreen role, meaning applications are continuously reviewed for current and future opportunities.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Ability to thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative across different stacks and teams to ensure successful shipping.

Nice To Haves

  • Experience training or supporting large-scale language models with tens of billions of parameters or more.
  • Experience working with reinforcement learning workloads (e.g., PPO, DPO, RLHF, or reward modeling).
  • Background in high-performance or reliability engineering — distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, OpenTelemetry).
  • Contributions to large-scale ML research or infrastructure, open-source frameworks, or internal performance optimization efforts.

Responsibilities

  • Design, build, and optimize the infrastructure for large-scale reinforcement learning and post-training workloads.
  • Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
  • Develop shared monitoring and observability tools for high uptime, debuggability, and reproducibility of RL systems.
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
  • Build evaluation and benchmarking infrastructure to measure model progress on helpfulness, safety, and factuality.
  • Publish and share learnings through internal documentation, open-source libraries, or technical reports.

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service