Research Engineer, Benchmarks

HUDSingapore, CA
Onsite

About The Position

HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25. We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs.

Requirements

  • Proficiency in Python, Docker, and Linux environments
  • Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc. - please link in your application
  • Strong understanding of what a “good benchmark” means and what makes one realistic, reliable, and useful
  • Experience working on environments and evals
  • Curiosity and ability to truly understand how workflows in various domains work

Nice To Haves

  • Detail-oriented and able to spot subtle inconsistencies or edge cases in tasks
  • Able to reason from first principles about task design, scoring, and failure modes
  • Thrive in unstructured problem spaces
  • Early-stage startup experience with ability to work independently in fast-paced environments
  • Strong communication skills for remote collaboration across time zones

Responsibilities

  • Own the design, implementation, and quality of HUD’s internal agent benchmarks
  • Work with subject-matter experts to define tasks and create domain-specific benchmarks that evaluate agents on realistic workflows
  • Build infrastructure to reliably run models and agents against benchmark tasks
  • Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
  • Validate whether benchmark performance correlates with real-world evals, customer needs, and lab expectations
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences

Benefits

  • Competitive compensation
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Equinox membership
  • 401k
  • Commuter benefits (US employees)
  • Unlimited access to tokens for ChatGPT, Claude Code, Cursor, etc.
  • Support for relocation and visas for strong full-time candidates to the US or Singapore
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service