RL Environments Engineer

Bespoke LabsMountain View, CA
Onsite

About The Position

Bespoke Labs is an applied AI research lab focused on curating data and RL environments for training and evaluating agents. They have developed significant resources like the Open Thoughts dataset and specialized models, and are positioned to lead in data and RL environment curation. This role is a delivery-focused engineering position responsible for building the systems that transform environment ideas into validated agentic coding tasks. The engineer will develop pipelines for programmatic environment production, design complex coding worlds for agent training, and continuously improve throughput (volume, quality, and efficiency). Success will be measured by the volume and quality of shipped environments, with a strong emphasis on prior experience in building and scaling such pipelines.

Requirements

  • Proven delivery of pipelines that produce RL environments or agentic tasks, with demonstrated personal contribution to volume.
  • Experience scaling task or environment creation into the hundreds or thousands through automation.
  • A track record of high throughput, with a focus on shipping products and reducing production costs per unit.
  • Strong software engineering fundamentals and Python fluency.
  • Real-world experience with production software, including large codebases, conventions, build systems, testing, devops, SRE, diagnosis, and RCA.
  • Understanding of the capabilities and limitations of frontier coding agents, including their common shortcuts.
  • Ability to build the infrastructure for scaled production: pipelines, automation, grading and verification systems, and sandboxed execution.
  • Experience running workloads at scale on GCP.
  • A tool-builder's instinct, with a focus on automating repetitive work and unblocking colleagues.

Nice To Haves

  • Hands-on experience with RL training systems, post-training, verifiers, or tool-use harnesses.
  • Background in developer tooling, CI/CD sandboxes, or code-execution infrastructure.
  • Experience in the review of task creation processes for benchmarks like Terminal Bench 3.0.

Responsibilities

  • Build environment-generation pipelines, owning systems for programmatic RL environment production including templating, automated grading, verification, and QA to enable scaled environment shipping.
  • Create complex coding worlds with high-fidelity environments that reflect real codebases, including conventions, dependencies, tooling, and technical debt.
  • Scale agentic task creation from small numbers to hundreds and thousands of validated tasks using automation.
  • Build internal tooling and infrastructure to identify and remove bottlenecks in environment production, increasing team throughput.
  • Own the full task lifecycle, including prompting, environment creation, grading, running frontier models, failure analysis, and iteration to ensure tasks are rigorous, fair, and difficult to game.
  • Defend quality at scale by catching reward hacking and grader loopholes, and establishing verification standards.
  • Direct coding agents heavily to accelerate environment building and validation, assessing their output and identifying subtle failures.

Benefits

  • Health coverage
  • Lunch
  • Flexible work arrangements
  • Opportunity to shape how the AI community evaluates and trains agents
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service