Product Designer, Evals & Prompts

AnthropicSan Francisco, CA
$305,000 - $385,000Hybrid

About The Position

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. The Product Prompt and Eval Design team is responsible for designing the system prompt that greets a new user, the tool descriptions that decide whether Claude searches, the instructions that keep a slide deck from tipping into slop, and the evals that test all of it. The goal is to ensure the model and the product align with user expectations, product strategy, and safety requirements across all surfaces and through every model launch. This role is a foundational member of the eval side of this work, focusing on building the evals that check prompts, the harness that runs them, and the tools that empower designers to perform this work independently. The position is part of the Product Prompt and Eval Design team within Product Design, collaborating daily with surface owners and engineers in each product team, and pairing with the prompt engineering team at model releases for each surface.

Requirements

  • Production-quality Python
  • Experience building and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, regression suites, and the underlying infrastructure for cross-model execution.
  • Experience building internal tools with user interfaces for non-coders.
  • Experience establishing test harnesses, sandboxing tool calls, and pinning settings for comparable run results.
  • Experience shipping prompts, or close collaboration with those who do, and understanding why a prompt effective on one model may fail on another.
  • Ability to read transcripts, not just analyze scores.

Nice To Haves

  • Experience working within a model-launch cycle.
  • A/B testing experience and the ability to link offline evaluations to online outcomes.
  • Front-end or notebook-to-app development experience, and informed opinions on how to make eval results easily understandable at a glance.
  • Experience translating product rubrics into training signals: graders, human-feedback questions, or preference pairs.
  • A genuine interest in Claude's behavior for its users, beyond just metric movement.

Responsibilities

  • Write and revise prompts for Claude's tools, features, and behaviors on product surfaces; test the surface, translate findings into prompt fixes, ship them, and confirm the prompt users receive is the intended one.
  • Build graders to validate prompt fixes and rerun them on subsequent models; convert designers' hand-run rubrics into automated evals, and analyze transcripts to identify what the eval missed.
  • Develop visual, low-code eval tools for designers to use without engineering assistance: assemble comparison sets from real transcripts, transform plain-English rubrics into graders, compare prompt variants across models side-by-side, and interpret results within the tool instead of a notebook.
  • Observe designers using these tools and iterate to simplify them.
  • Support model releases by testing each surface against the new model, writing prompt fixes and migrations, and creating prompts for features launching with it, ensuring surface owners have data-driven decisions.
  • Establish and scale the eval harness: build the test environment to exercise 50-100 tools with trustworthy settings, maintain green evals across models, and determine if regressions stem from the harness or the model.
  • Package elements that prompting cannot fix for training, along with the attached eval: create graders for precise behaviors and human-feedback questions/good-bad pairs for subjective aspects like writing quality.

Benefits

  • Competitive compensation
  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
  • Office space for collaboration
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service