Principal AI Quality Engineer

Nexaminds
Remote

About The Position

Nexaminds is looking for an AI Quality Engineer to lead the validation and quality strategy for AI-powered systems, including Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) applications, and agent-based workflows. The ideal candidate has deep experience in Quality Engineering, AI evaluation, and automated validation frameworks, with a strong understanding of how to measure, monitor, and improve the reliability of non-deterministic AI systems. This role focuses on building enterprise-grade AI validation capabilities, designing automated evaluation pipelines, implementing AI guardrails, and partnering closely with engineering, AI/ML, and data governance teams to ensure AI-powered solutions are accurate, safe, and production-ready.

Requirements

  • 7+ years of experience in Quality Engineering, Software Development, or a related technical discipline, including recent hands-on experience with AI/ML or LLM-based systems.
  • Experience designing validation or evaluation frameworks for non-deterministic or probabilistic systems such as Large Language Models (LLMs) or Machine Learning applications.
  • Strong understanding of Retrieval-Augmented Generation (RAG), agentic workflows, and LLM evaluation concepts, including hallucination detection, groundedness, relevance, and consistency.
  • Hands-on experience with Large Language Models (LLMs) and prompt engineering, including prompt design, optimization, and regression validation.
  • Practical experience using Claude Code, Claude SDK, or similar AI-assisted development platforms and agentic coding tools.
  • Proficiency with Node.js for building automation tools, validation pipelines, and engineering utilities.
  • Experience working with MongoDB to store, manage, and analyze evaluation results, validation data, or AI output logs.
  • Experience integrating automated validation processes into CI/CD pipelines.
  • Experience designing semantic validation approaches that evaluate meaning, relevance, and contextual accuracy beyond traditional software testing.
  • Strong understanding of AI quality metrics, confidence scoring, output consistency, and production monitoring.
  • Solid knowledge of data validation principles, including schema validation, business rule compliance, and data consistency.
  • Excellent cross-functional collaboration skills, with experience partnering across Quality Engineering, AI/ML Engineering, Software Engineering, and Data Governance teams.

Nice To Haves

  • Experience with prompt regression testing tools or frameworks.
  • Familiarity with AI governance, fairness, or bias-detection practices.
  • Experience with tools such as Playwright, RestSharp, or similar UI/API test automation.
  • Prior experience standing up a validation or evaluation function from scratch.
  • Exposure to confidence scoring or guardrail systems for production AI.
  • Experience designing or orchestrating multi-agent systems in production environments

Responsibilities

  • Define and lead the AI validation strategy for LLMs, RAG systems, and AI agents.
  • Build automated validation pipelines integrated into CI/CD.
  • Establish evaluation frameworks covering correctness, relevance, groundedness, consistency, and hallucination rate.
  • Design and maintain prompt regression testing to catch silent quality degradation after model or prompt changes.
  • Introduce semantic validation techniques that go beyond traditional QA (evaluating meaning, not just structure).
  • Lead production monitoring of AI output quality and define acceptable output ranges for non-deterministic systems.
  • Partner with engineering teams to validate AI-generated code, test cases, and AI-assisted decision workflows.
  • Introduce AI guardrails and confidence scoring to flag unsafe or low-confidence outputs.
  • Build validation checkpoints across multi-step agent pipelines.
  • Define and track quality metrics/KPIs (defect reduction, output consistency, prompt regression stability, adoption confidence).

Benefits

  • Stock options
  • Remote work options
  • Flexible working hours
  • Benefits above the law
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service