LLM Model Response Evaluation

Lifted, an Upwork Company™Maryland City, MD
Remote

About The Position

Our client, a global technology company that helps businesses build, train, and manage AI systems, is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria. This role involves evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities. The work may encompass text, images, audio, video, HTML widgets, PDFs, or combinations thereof. The role is domain-agnostic, covering topics across arts, culture, history, science, engineering, and more. Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions. Detailed project guidelines will be provided within the evaluation platform for each task.

Requirements

  • 3+ years of hands-on experience in LLM / GenAI data evaluation.
  • Master's or PhD required (PhD candidates strongly preferred).
  • Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
  • Comfortable evaluating content across multiple modalities (text, images, audio, video, HTML widgets, PDFs).

Responsibilities

  • Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
  • Researching unfamiliar topics using trusted sources before making evaluation decisions.
  • Making well-supported judgments based on research and project guidelines.

Benefits

  • Flexible and remote work
  • Variable workload: Accept or decline tasks based on your availability
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service