AI Evaluation Engineer

DialpadVancouver, BC
CA$115,500 - CA$132,750Hybrid

About The Position

As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions. This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.

Requirements

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.

Responsibilities

  • Design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • Build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
  • Co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • Create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
  • Develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
  • Investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
  • Collaborate with cross-functional teams, including applied science, engineering, and Product QA.

Benefits

  • Competitive salary
  • comprehensive benefits
  • real opportunities for growth
  • cutting-edge AI tools
  • a robust training program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service