Assessment Design Lead

SpeakSan Francisco, CA
$100,000 - $140,000

About The Position

Speak cares deeply about learners actually learning and improving with Speak. We have a dedicated Proficiency team to own how we measure learning efficacy and speaking proficiency, from unit-level mastery checks to standalone proficiency tests to onboarding placement, all in service of understanding users’ proficiency levels and learning gains in an accurate, transparent and actionable manner. Because everything happens remotely and asynchronously in the app, keeping scores fair, stable over time, and resistant to gaming is both hard and genuinely interesting. We're looking for an Assessment Design Lead: the person who defines what we measure, why, how, and signs off on whether our assessments actually measure the right thing. You will be staffed on the Proficiency team and report to the Head of Learning Design and Curriculum. If solving reliable, at-scale speaking assessment excites you and you want to directly shape Speak's efficacy story, we'd love to hear from you.

Requirements

  • 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) — not just academic theory.
  • Can explain reliability and validity in plain language and knows how to catch a test that's measuring the wrong construct.
  • Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."
  • Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) — essential for a speech-based test serving learners across dozens of native languages.
  • Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.
  • Comfortable running or interpreting the statistics behind a rubric or rater system — inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts — enough to know whether a scoring system is actually reliable, not just plausible.
  • Comfortable being the sign-off authority on content validity — makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.
  • Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite — distinct from spec'ing work for the ML Engineer to build.
  • Comfortable operating in a 0-to-1 environment.
  • Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup.

Nice To Haves

  • Speech/pronunciation science background — can own pronunciation frameworks and L2-specific error taxonomy directly
  • Familiarity with adaptive testing or IRT-adjacent concepts (even if not the primary psychometrician)
  • Experience at a large-scale language testing organization or similar high-rigor assessment environment
  • Experience thriving in an EdTech startup environment, especially in a newly forming team or 0-to-1 mandate
  • Advanced degree (Master's or PhD) in psychometrics, measurement, applied linguistics, SLA, or a related quantitative field; track record of shipped assessment work still matters more
  • Has authored technical/validity reports or published assessment research

Responsibilities

  • Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) — with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve.
  • Keep constructs (fluency, pronunciation, grammar, task achievement) clearly separated and each aligned to CEFR or a comparable speaking proficiency standard, so no single assessment conflates domains it wasn't designed to measure.
  • Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.
  • Sign off on content validity for every assessment that ships.
  • Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups.
  • Making sure scores stay comparable as the assessment evolves and stay meaningful against attempts to game an unproctored test.
  • Design and run the validity evidence plan — so validity is built into the process rather than checked only after launch.
  • This includes concurrent/criterion studies benchmarking Speak’s assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.
  • Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation.
  • You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.

Benefits

  • Opportunity to travel
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service