Principal Research Scientist - Evaluations

CanvaSan Francisco, CA
$310,000 - $420,000Hybrid

About The Position

Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships. We are looking for a Principal Research Scientist who defines what evaluation needs to become as the space gets harder, rather than running the playbook we already have. You will own how Canva evaluates generative quality across the whole of Canva Research, including problems we have not framed yet: new modalities, evaluation that reflects real differences between content types, user segments and markets, and a much tighter link between what our metrics say and what users and the business actually experience. This is a Canva-wide craft leadership role, setting direction across our research groups in Australia, Europe, the US and China. You will be the person others come to when the numbers and the eyes disagree.

Requirements

  • A track record of defining evaluation frameworks and standards from the ground up rather than operating inside someone else's, ideally at an organisation pushing the frontier of GenAI evaluation, and measurement systems that changed how a team made decisions rather than papers about metrics
  • Experience linking evaluation metrics to downstream business or user outcomes, and diagnosing why they diverge
  • Experience turning subjective human judgement into reliable, objective evaluation signal through rubric design, human data pipelines and model training
  • Strong grounding in multimodal generative models (diffusion, transformers, VLMs and MLLMs) and their architectures, deep enough to know where evaluation will break
  • Experience with reward modelling, preference learning, or alignment methods involving human feedback
  • Experience setting technical direction across multiple teams or a wide specialty area in a globally distributed organisation, with the instinct to find the gap nobody owns and close it without waiting for a mandate

Nice To Haves

  • Research background in human perception, psychophysics, aesthetics or HCI
  • Experience evaluating for harm, bias and safety alongside quality
  • Publication record in evaluation, alignment or generative modelling
  • Background or genuine interest in visual arts and graphic design

Responsibilities

  • Establish the evaluation gates that inform launch decisions, and be accountable for the judgement calls when the signal is ambiguous
  • Diagnose anomalous evaluation results during production training runs, separate model regressions from infrastructure artefacts, and communicate the answer clearly and quickly
  • Partner with Design Generation, Foundation Models and Agents teams so evaluation shapes their training and inference, rather than reporting on it after the fact
  • Work directly with designers, creators and product teams to turn subjective creative judgement into measurable criteria
  • Mentor senior research scientists and engineers, and raise the bar for evaluation craft across the group
  • Represent Canva's evaluation vision and practice to senior leadership and, where valuable, to the broader industry community

Benefits

  • Equity packages
  • Health benefits plans
  • 401(k) retirement plan with company contribution
  • Inclusive parental leave policy
  • An annual Vibe & Thrive allowance
  • Flexible leave options
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service