Deeptune is an RL environments lab within Mercor that builds training gyms for AI agents: high-fidelity simulations where AI learns to perform real-world tasks through reinforcement learning. We work with the leading AI labs to help them train the next generation of agentic models, and our environments have already contributed to recent breakthroughs in computer use, code generation, and multi-step task completion. This role requires high ownership and abstraction: setting direction, driving outcomes, and staying hands-on while leading. You'll also build direct partnerships with leading AI labs and enterprises, and contribute hands-on to improving frontier model quality through data, evaluation, and systems — work that sits squarely in the areas the field agrees matter most right now: RL with verifiable rewards, rubric-based reward modeling for subjective and agentic domains, and the pipelines that turn raw human demonstrations into training-ready environments.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed