Research Engineer - Evaluations

LumaRedwood City, CA

About The Position

As a Research Engineer on the Evaluations team, you will own the infrastructure responsible for assessing the improvement of Luma's models. Your role will involve building pipelines, metrics, and automated systems to connect model output with measurement and improvement, ensuring that every model decision is based on rigorous and consistent data. This position spans research, engineering, and product, requiring expertise in both systems and machine learning. You will be responsible for transforming human evaluation frameworks into automated ones and integrating the results directly into the training process. This role is ideal for someone who enjoys owning both the ML aspects and the surrounding infrastructure. If your interest lies solely in model development without involvement in measurement pipelines, this position may not be suitable.

Requirements

  • 5+ years building ML evaluation systems, model pipelines, or large-scale infrastructure.
  • Master's or PhD in Computer Science, Machine Learning, or a related field, or equivalent industry experience.
  • Hands-on experience with visual data (image and/or video) in evaluation, modeling, or data preparation.
  • Proficiency in Python and an ML framework (PyTorch, JAX, or TensorFlow).
  • Strong ML background with generative models (diffusion, LLMs, multimodal architectures).
  • Strong software engineering skills: CI/CD, testing, data pipelines, distributed systems.
  • Familiarity with human-in-the-loop evaluation and methods for scaling it through automation.

Nice To Haves

  • Experience with reinforcement learning or reward modeling.
  • Prior work on perceptual metrics, multimodal benchmarks, or retrieval-based evaluation.
  • Background in large-scale model training or evaluation infrastructure.
  • Familiarity with creative media workflows (film, VFX, animation, digital art).
  • Contributions to open-source evaluation libraries or benchmarks.

Responsibilities

  • Design and build scalable pipelines for automated evaluation of generative models across image, video, text, and audio.
  • Develop metrics and evaluation models that capture fidelity, coherence, temporal consistency, and alignment with human intent.
  • Integrate evaluation signals into training loops, including reinforcement learning and reward modeling, to continuously improve models.
  • Build the infrastructure for large-scale regression testing, benchmarking, and monitoring of multimodal models.
  • Collaborate with researchers conducting human studies to translate their frameworks into automated or semi-automated systems.
  • Maintain dashboards, reporting, and alerting systems to surface evaluation results to relevant teams.

Benefits

  • Equal opportunity employer
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service