AI Voice Evaluation Specialist

Innodata•Florida, DE

About The Position

Innodata is a global data engineering company focused on enabling the responsible advancement of artificial intelligence. We provide the data, evaluation frameworks, and human expertise needed to build trustworthy AI systems at scale. We offer a range of solutions, platforms, and services for Generative AI/AI builders and adopters, leveraging our 36+ year legacy of delivering high-quality data and outstanding customer outcomes. This role involves evaluating the performance of AI models through real-time, voice-based conversations. Specialists will interact with two different AI models using the same assigned scenario, compare their responses, and provide structured evaluations based on defined quality criteria. The primary goal is to ensure a fair and consistent comparison between models and identify which model delivers a stronger overall conversational experience.

Requirements

  • Bachelors Degree
  • Strong attention to detail and ability to notice subtle differences in conversational quality.
  • Excellent listening and comprehension skills.
  • Strong written communication skills, particularly the ability to explain observations clearly and objectively.
  • Ability to follow detailed evaluation guidelines consistently.
  • Comfort speaking naturally and roleplaying different conversational scenarios.
  • Ability to compare two interactions fairly without allowing personal preferences to influence the evaluation.
  • Strong critical-thinking and analytical skills.
  • Reliability and consistency when completing structured evaluation tasks.
  • Familiarity with AI assistants, voice-based AI, or conversational systems.

Responsibilities

  • Review an assigned scenario and roleplay a natural conversation with two different AI models.
  • Conduct comparable conversations with Model A and Model B, keeping the scenario, conversational approach, and overall interaction as consistent as possible.
  • Maintain approximately the same number of conversational turns with each model to support a fair comparison.
  • Record voice during each interaction and pay close attention to the quality and behavior of the model's responses.
  • Evaluate each model across five defined evaluation dimensions.
  • Identify and categorize relevant error clusters that may occur in the audio or conversation.
  • Compare the performance of Model A and Model B based on the evaluation criteria and observations.
  • Select the model that demonstrated the stronger overall performance according to the defined evaluation framework.
  • Write a detailed rationale explaining the final preference, referencing specific examples and observations from both conversations.
  • Apply evaluation guidelines consistently across different scenarios and model interactions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service