AI Voice Evaluation Specialist

Innodata Inc.Indianapolis, MA
$20 - $27

About The Position

Innodata is a global data engineering company focused on enabling the responsible advancement of artificial intelligence. We provide data, evaluation frameworks, and human expertise to build trustworthy AI systems. We offer a range of solutions, platforms, and services for Generative AI / AI builders and adopters, leveraging our 36+ year legacy of delivering high-quality data and outstanding outcomes. This role involves evaluating the performance of AI models through voice-based conversations. Specialists will interact with two different AI models using the same assigned scenario, compare their responses, and provide structured evaluations based on defined quality criteria. The primary goal is to ensure a fair and consistent comparison between models and identify which model offers a superior conversational experience.

Requirements

  • Bachelors Degree
  • Strong attention to detail and ability to notice subtle differences in conversational quality.
  • Excellent listening and comprehension skills.
  • Strong written communication skills, particularly the ability to explain observations clearly and objectively.
  • Ability to follow detailed evaluation guidelines consistently.
  • Comfort speaking naturally and roleplaying different conversational scenarios.
  • Ability to compare two interactions fairly without allowing personal preferences to influence the evaluation.
  • Strong critical-thinking and analytical skills.
  • Reliability and consistency when completing structured evaluation tasks.
  • Familiarity with AI assistants, voice-based AI, or conversational systems.

Responsibilities

  • Review assigned scenarios and roleplay natural conversations with two different AI models.
  • Conduct comparable conversations with Model A and Model B, maintaining consistency in scenario, conversational approach, and interaction.
  • Maintain approximately the same number of conversational turns with each model for a fair comparison.
  • Record voice during each interaction, paying close attention to the quality and behavior of model responses.
  • Evaluate each model across five defined evaluation dimensions.
  • Identify and categorize relevant error clusters in audio or conversation.
  • Compare the performance of Model A and Model B based on evaluation criteria and observations.
  • Select the model that demonstrated stronger overall performance according to the defined evaluation framework.
  • Write a detailed rationale explaining the final preference, referencing specific examples and observations from both conversations.
  • Apply evaluation guidelines consistently across different scenarios and model interactions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service