About The Position

As Luma's Qualitative Evaluation Engineer, you will be responsible for defining and implementing frameworks to assess the quality of generative models beyond quantitative metrics. This role involves translating abstract qualities like believability, identity retention, and scene coherence into clear, testable criteria. You will build evaluative systems that align with human perception and creative intent, working closely with researchers and technical artists to steer model development. This position is ideal for someone who thrives on defining fuzzy qualities and creating actionable insights from nuanced judgments, rather than focusing solely on quantitative dashboards.

Requirements

  • 5+ years in product evaluation, UX research, model testing, or similar structured qualitative assessment.
  • Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or a related field.
  • Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX).
  • Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms.
  • Excellent written communication and the ability to synthesize nuanced judgment into actionable insight.
  • Comfort working across engineers, researchers, and creatives.

Nice To Haves

  • Background in motion, visual effects, or storytelling pipelines.
  • Experience evaluating AI-generated media (video, images, 3D).
  • Prior work building internal tools for qualitative data collection or scoring.
  • Familiarity with prompt engineering and reference-based inputs.

Responsibilities

  • Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality.
  • Build and maintain qualitative evaluation frameworks that are scalable and reusable.
  • Translate high-level product goals into concrete evaluative criteria.
  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations.
  • Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX.
  • Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service