Quality Lead, Agentic AI Workflow Evaluation

InnodataSan Jose, CA
Onsite

About The Position

Innodata is seeking a Quality Lead for a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. Reviewers will work through ambiguous, multi-step scenarios in isolated test environments, assessing AI agent performance regarding safety, user intent, consent, and overall robustness. This senior individual contributor role is accountable for the quality of the output. The Quality Lead will set standards for reviewers, own the audit sample, run calibration, maintain the rubric, and train reviewers. This role also serves as a deputy to the Engagement Manager. The quality approach is not fully defined and will be built in partnership with the customer's quality leads, or adapted from existing frameworks with recommendations for improvement.

Requirements

  • Bachelor's degree or equivalent practical experience
  • 4+ years in quality assurance, quality management, or senior review work within annotation, evaluation, trust and safety, or a similarly judgment-intensive domain
  • Direct experience owning a quality function: you designed the audit, not just executed someone else’s
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review
  • Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments, and common failure modes
  • Demonstrated ability to run calibration with peers — including holding a position under disagreement and changing it when the argument is better
  • Strong written communication; able to document a scoring standard clearly enough that a reviewer can apply it and an auditor can check it
  • Comfortable in spreadsheets and in a dashboarding tool, with enough Python or SQL to pull and slice your own data
  • Experience training or onboarding reviewers into rubric-based work

Responsibilities

  • Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported.
  • Build the quality system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on data.
  • Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes.
  • Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning.
  • Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer.
  • Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready.
  • Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working.
  • Report quality trends to the Engagement Manager and, alongside them, to the customer.
  • Deputize for the Engagement Manager on delivery operations during absences.
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service