Human Evaluation Researcher

Nuance LabsSeattle, WA
$160,000 - $190,000Onsite

About The Position

Nuance Labs is developing advanced AI avatars with emotional intelligence, aiming for photorealistic, real-time interaction that mimics human conversation. The company is composed of experienced researchers and engineers from top institutions and tech companies, backed by prominent investors. They are focused on overcoming the limitations of current conversational AI, which are often slow and lack emotional depth, by building a full-duplex system that perceives and expresses emotion in real-time. This role is crucial for evaluating the human-like qualities of these avatars, as automated metrics cannot capture nuances like sincerity, warmth, or trustworthiness.

Requirements

  • 5+ years designing and running human-subjects research in industry or academia (UX research, HCI, experimental psychology, behavioral science, or a related field).
  • Ability to provide examples of study designs, particularly those that achieved human convergence on ambiguous judgments (tone, emotion, quality, trust), including the rubrics, anchors, and protocols used.
  • Strong grounding in both qualitative methods (interviews, ethnography, contextual inquiry) and quantitative methods (survey and psychometric design, experimental design, statistics for rating and pairwise-comparison data).
  • Fluency with agreement and reliability metrics (e.g., Cohen's kappa, Krippendorff's alpha) and methods to improve them.
  • A bias toward executing scrappy, sound studies quickly and the ability to explain findings clearly to ML researchers.

Nice To Haves

  • Experience evaluating generative AI, such as avatars, digital humans, speech or video generation, conversational agents, or emotion expression and recognition.
  • A background in perceptual science or psychophysics, focusing on human perception of faces, voices, motion, and emotion (MS/PhD in a related field is a plus).
  • Familiarity with large-scale human evaluation, including crowdsourcing platforms, annotation tooling, and golden datasets.
  • Proficiency in statistics and scripting (Python or R) for data analysis.

Responsibilities

  • Design and run qualitative and quantitative studies of AI avatars, including side-by-side comparisons, controlled rating experiments, in-depth interviews, think-alouds, diary studies, and longitudinal panels.
  • Develop instruments like rubrics, anchored scales, and annotation guidelines to translate ambiguous human judgments into measurable signals and improve inter-rater agreement.
  • Utilize ethnographic techniques such as observation of live conversations, contextual inquiry, and field work to understand user experiences with AI avatars.
  • Build the human evaluation pipeline, encompassing participant panels, rater training and calibration, tooling, and a cadence for evaluations aligned with model releases.
  • Calibrate automated and model-based metrics against human judgment to determine their reliability.

Benefits

  • Health insurance with an HDHP option and ~$2,000 in annual HSA contributions.
  • 15 days of PTO.
  • 10 paid public holidays.
  • Office closure for a full week at year-end.
  • Provided lunch, drinks, and snacks daily.
  • Boba Tuesdays and Thursdays.
  • Commuter benefits for parking and transportation (up to $340/month pre-tax).
  • 401(k) with a 4% match (100% of the first 1% and 50% of the next 5% contributions).
  • Meaningful equity.
  • Visa sponsorship (O-1, H-1B, green card, etc.) from day one.
  • AI-native tooling with unlimited tokens.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service