About The Position

At Pencil, we are driving innovation in advertising technology through our state-of-the-art SaaS product, which harnesses Generative AI to redefine content creation. Our mission is to make AI the default in advertising without replacing creative people. To achieve this, we need to make sure that our technology isn't just in the hands of big brands - we need to try to help small businesses and creative individuals too. The role: Pencil's agents produce advertising creative at scale - video, adaptation and multiple formats, orchestrated by our agentic system, Scribble. The product only works if the output is right first time. This role owns measuring quality and improving it. You'll own evals end to end. Internally, that means the systems that measure whether agents, skills and workflows produce work a client would ship. Externally, it means how we demonstrate quality to customers and the market: the evidence behind client RFP responses, creative enablement packages, and adevals.ai , our public evals site. Measuring creative quality is hard because there is no single right answer. Your job is to build reliable measurement anyway: rubrics that define what good means, calibration between human reviewers and automated judges, and a golden dataset that stays representative as the product and client base evolve. You'll then make those metrics central to how the business operates and sells.

Requirements

  • A quality, evals or ML measurement problem you owned as a product, with its own roadmap and users.
  • Measurement you shipped in a subjective domain: how you defined quality, and how you kept the signal reliable as it scaled.
  • Direct experience with the mechanics: golden datasets, rubric design, judge calibration, inter-rater reliability. Where they failed and what you did.
  • A product decision you derived from evidenced customer need in a domain you had to learn.
  • A metric you owned that changed decisions outside your own team.
  • Time spent in front of clients: presenting or defending methodology to buyers.

Nice To Haves

  • If your instinct is to start with a solution, this isn't the role.

Responsibilities

  • Own the charter for Evals & Quality: the durable problem, success metrics, and an explicit in/out scope list.
  • Define the core quality metrics - first-pass-right, coverage, regression detection - and drive them up.
  • Own the judging methodology: rubrics for creative quality, and the calibration process that keeps automated judges and human raters aligned over time.
  • Design and run the human QC mechanism: where humans review output, how their judgements feed the automated layer, and how it scales.
  • Own the golden dataset: curation, coverage across formats and client contexts, and versioning.
  • Build eval loops that other teams use themselves. A team shipping a new video skill should be able to set up evals without your involvement.
  • Own adevals.ai as a product, including its design quality. The site needs to meet a high design bar to be credible.
  • Provide the evidence layer for commercial work: client RFP responses and creative enablement packages.
  • Partner with Agent Architects on client-specific quality bars, delivered through configuration and rubrics rather than custom eval code.
  • Run the full product lifecycle: Discovery through post-launch audit, with a written hypothesis before every ship and a review 14 days after launch.

Benefits

  • 25 days PTO plus public holidays, although we operate a Flexible Time Off scheme
  • Health insurance / private medical cover
  • Monthly stipend towards wellness, fitness, and learning and development
  • Remote - work from anywhere in your home country
  • Enhanced parental leave policies, whether you become a parent through birth, adoption or surrogacy
  • Access to our Pencil office in The Shard, London for our UK employees
  • Flexible working hours
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service