AI Evaluation Engineer (QA)

Appnovation TechnologiesToronto, ON

About The Position

As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale — from small human-UAT batches up to millions of automated evals — statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time. We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done.

Requirements

  • Bachelor’s Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.
  • You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
  • You understand lean thinking
  • You set high standards for code quality, performance/scalability and security and seek continuous improvement
  • You have solid analytical, problem solving and decision-making skills
  • You have customer first mindset and a devotion to customer service
  • You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
  • You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
  • You are responsive and thrive in a fast-paced diverse high-performance environment with rapidly changing business needs
  • You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
  • You have prior experience in consulting

Nice To Haves

  • Prior experience and connections in the Life Sciences industry is preferred.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Benefits

  • Accommodations are available upon request throughout the recruitment process.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service