HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25. We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed