Member of Technical Staff, Frontier Evals

IntelligenceSan Francisco, CA
Onsite

About The Position

Intelligence is the product lab company behind DesignArena - 5.5M+ users in 8 months, leaderboard referenced by Andrew Ng, Elon Musk, Demis Hassabis, and more. The team is incredibly talent-dense (11 from Harvard + Berkeley), backed by Tier 1 VC Index Ventures, YC, SV Angel, Lenny Rachitsky, Paul Graham, Dylan Field, and one of the fastest-growing seed-stage startups in SF. Behind closed doors, we see what the models can do six months before the world does. We are trusted by the best frontier model providers like OpenAI to rigorously evaluate the capabilities of state-of-the-art multimodal models across design, web dev, game dev, image, video, audio, slide generation, and more, through the large-scale platforms that we’ve built. Design Arena, our flagship product, is the most referenced benchmark for AI-generated visuals, and is powered by over 5.3M+ authentic users across 192 countries. Prediction Arena was the first time models traded autonomously with real cash on real-time, real-world events. Social Arena tested whether AI models can effectively grow and engage audiences on X by having them operate as independent social media agents. You'll define how frontier AI models are measured. You'll design new benchmarks, run experiments, analyze model behavior, and build evaluation methodologies that become trusted signals for the industry. Your work will shape our public leaderboards and the evaluation tools we share with frontier labs. Here’s an example of a piece of industry-leading work done in this field. This is a SOTA STS benchmark, advised by OpenAI: https://audioarena.ai/. Email us for the pre-print.

Requirements

  • Strong STEM background. You studied Computer Science, Data Science, Statistics, Math, Engineering, Physics, or a related field.
  • Deep curiosity about frontier AI models. You're excited by understanding model behavior, discovering areas of failure, and building better ways to evaluate models.
  • Genuine thirst and intellectual to be on the frontier of AI development.
  • Fearlessness to roll up your sleeves, get your hands dirty, and do real work.

Responsibilities

  • Design genuinely hard and useful evaluations that measure frontier model performance on real-world tasks, and that become industry-leading gold-standards
  • Investigate model failures and identify what they reveal about emerging capabilities
  • Publish research, technical reports, and analyses that shape how frontier models are evaluated

Benefits

  • Competitive salary + meaningful equity.
  • Sponsor visas and handle relocation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service