RL Environments Engineer

Bespoke LabsMountain View, CA
$250,000 - $300,000Onsite

About The Position

Bespoke Labs is an applied AI research lab focused on curating data and RL environments for training and evaluating agents. They have developed significant datasets and models, and are positioned to capture a large market share in data and RL environment curation. This role is a delivery-focused position for an engineer who has experience building machinery to turn environment ideas into a large number of validated agentic coding tasks. The engineer will be responsible for building pipelines that mass-produce environments, designing complex coding worlds for agent training, and increasing throughput (more environments, higher quality, less manual work). Success will be measured by the volume and quality of environments shipped, not by publications. The most important qualification is prior experience in standing up an environment-generation pipeline and scaling agentic task creation.

Requirements

  • A record of shipped volume in building agentic coding tasks or environments, with the ability to demonstrate personal contribution and production cost.
  • Experience scaling output through automation rather than manual labor.
  • Strong software engineering fundamentals and fluency in multiple production-ready languages.
  • Real experience with production software, including large codebases, build systems, testing, deployment, on-call responsibilities, and root cause analysis.
  • An adversarial mindset, capable of identifying and fixing potential model cheating in graders.
  • A clear understanding of the capabilities and limitations of frontier coding agents, including their tendency to cut corners.
  • Ownership mentality, with the ability to build, debug, and ship without significant supervision.

Nice To Haves

  • Experience with RL training systems, post-training, verifiers, or tool-use harnesses.
  • Background in developer tooling, CI/CD, sandboxes, or code execution infrastructure.
  • Experience building large-scale automated test generation, fuzzing harnesses, or benchmark suites.
  • Contributions to public agentic benchmarks like Terminal-Bench.
  • Open-source contributions that are depended upon by others.

Responsibilities

  • Build environment-generation pipelines, owning systems for programmatic RL environment production including templating, automated grading, verification, and QA to enable scaled environment shipping.
  • Create complex coding worlds with high-fidelity environments that reflect real codebases, including their conventions, dependencies, tooling, and technical debt.
  • Scale agentic task creation from small numbers to hundreds and thousands, utilizing automation for efficiency.
  • Build internal tooling and infrastructure to identify and remove bottlenecks in environment production, thereby increasing team throughput.
  • Own the full task lifecycle, from prompt and environment design to grader development, frontier model execution, failure analysis, and iteration, ensuring tasks are rigorous, fair, and difficult to game.
  • Defend quality at scale by identifying and preventing reward hacking and grader loopholes, and establishing verification standards as volume increases.
  • Direct frontier coding agents heavily to accelerate environment building and validation, assessing their output and identifying subtle failures.
  • Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.

Benefits

  • Health, dental, and vision coverage
  • 401(k)
  • Daily onsite lunch provided
  • Visa sponsorship and relocation support available
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service