Member of Technical Staff, Evals

Handshake•Mountain View, CA
•$200,000 - $350,000•Hybrid

About The Position

Handshake AI works directly with frontier labs on their most consequential data, evaluation, and post-training challenges, building the systems that turn expert human knowledge into the data and evaluations that make frontier models better. This role is for someone who is deeply passionate about building software with agents and creating the evaluations that reveal where those agents truly succeed and fail. You will design and publish coding benchmarks that meaningfully challenge state-of-the-art agents. You will work with AI researchers, software engineers, and domain experts to turn difficult, real-world software tasks into high-signal evaluation environments, datasets, verifiers, and feedback systems. The work will help shape both how the frontier evaluates coding agents today and where the field goes next. Early members of the team will have unusual influence over our technical direction, operating culture, and the open-source benchmarks, software, and research products we build. We care more about demonstrated technical depth, judgment, and a builder's mindset than a specific title, degree, or career path.

Requirements

  • Deep enthusiasm for agentic software development, with clear evidence that you actively build, experiment with, or think seriously about coding agents.
  • Strong software engineering skills and the ability to write clean, reliable, maintainable code.
  • Strong Python skills and comfort working with modern ML tooling, evaluation infrastructure, and data workflows.
  • Sound experimental judgment: you can form hypotheses, choose meaningful metrics, diagnose failures, and distinguish genuine capability improvement from evaluation artifacts.
  • Experience designing systems—not only implementing specifications—including the ability to make tradeoffs around validity, quality, scale, reliability, and reuse.
  • Comfort operating in an ambiguous, fast-moving environment with substantial ownership.
  • Collaborative, low-ego communication and the ability to work effectively with researchers, engineers, domain experts, and customers.
  • Experience working on a widely used coding-AI benchmark or evaluation suite.
  • Published research on AI for coding or code generation at a leading venue, such as NeurIPS, ICML, ICLR, or COLM.
  • Software engineering experience at a top technology company, or a strong public GitHub profile, paired with deep knowledge of agentic AI and a clear passion for building software with agents.

Nice To Haves

  • Building or maintaining coding benchmarks, coding-agent environments, repository-level evaluation suites, or open-source developer tools.
  • Developing automated graders, test harnesses, programmatic verifiers, reward models, or reinforcement-learning environments for software tasks.
  • Researching code generation, program synthesis, AI agents, reinforcement learning, post-training, or software-engineering productivity.
  • Experience with post-training methods such as reinforcement learning, RLHF, preference optimization, supervised fine-tuning, or reward modeling.
  • Strong public contributions through GitHub, papers, benchmarks, technical writing, or developer communities.
  • Experience turning research prototypes or repeated customer work into robust, reusable products or platforms.

Responsibilities

  • Design, build, and publish coding benchmarks that measure meaningful progress in frontier coding agents.
  • Create realistic, difficult software-engineering tasks, repositories, environments, and test harnesses that expose agent capabilities and failure modes.
  • Develop reliable verifiers, graders, reward signals, and evaluation methodology for agentic software development.
  • Research how to make coding-agent evaluations representative, difficult, robust, and resistant to shortcutting or benchmark contamination.
  • Analyze coding-agent behavior and trajectories to understand where agents fail, what feedback is useful, and which capabilities matter next.
  • Partner directly with AI researchers, software engineers, and expert contributors to develop high-signal tasks, data, and evaluation methods.
  • Run fast, rigorous iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and turn learnings into the next benchmark or system.
  • Identify repeatable patterns across engagements and productize them into reusable software, benchmarks, datasets, and platforms.
  • Raise the technical bar through strong design judgment, clear communication, code quality, and mentorship.
  • Contribute to the field through open benchmarks, open-source tools, research, and technical writing where it creates leverage.

Benefits

  • Equity in a fast-growing company
  • 401(k) match
  • Competitive compensation
  • Financial coaching
  • Paid parental leave
  • Fertility benefits
  • Parental coaching
  • Medical, dental, and vision
  • Mental health support
  • $500 wellness stipend
  • $2,000 learning stipend
  • Ongoing development
  • Commuting support
  • Free lunch
  • Gym in our SF office
  • Flexible PTO
  • 15 holidays + 2 flex days
  • Team outings
  • Referral bonuses
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service