Humanity is in a virtuous cycle: human insight improves AI, and better AI expands what people can do. Sustaining it depends on the one input that can't be automated: expert human judgment. At Pareto, we build the platform that turns that judgment into the data, evals, and RL environments frontier models learn from. We work with leading frontier labs like Anthropic and GDM, and we give skilled people everywhere a way to shape the future of AI and share in what it creates. This RL environment and human-data infrastructure is already in production. Our job now is to scale it. You'll own the RL environments frontier labs train on, end to end. Scope the problem with the requester, build the image and the tools inside it, write the graders that score it, ship it into the customer's platform, and keep it healthy once it's running. You sit between Pareto's engineering team and the researchers at the labs we work with, close enough to both that you can tell when a training goal and a buildable spec have drifted apart. Nobody will hand you a finished spec. You'll get a research problem, define what gets built, and stay with it after it lands. In your first year, good looks like environments that ship faster than the last one did, because you invested in the build and release path instead of hand-rolling each delivery. What you build becomes training signal. That's the reason the ownership runs all the way through production.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed