Manager, AI Engineering (Tester )

MastercardO'fallon, MO
$140,000 - $231,000Onsite

About The Position

Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready. This role will lead AI quality engineering efforts — defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates — while fostering a culture of rigorous, first-class AI testing across the program.

Requirements

  • Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
  • Demonstrated expertise testing LLM and Gen AI systems — including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
  • Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, or equivalent platforms.
  • Strong understanding of Gen AI failure modes — hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures — and proven methods to surface and document them systematically.
  • Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.
  • Working knowledge of LLM ecosystems — OpenAI, Anthropic, Hugging Face, LangChain/LangGraph — sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
  • Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
  • Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production.
  • Strong analytical, communication, and stakeholder management skills — with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences.

Responsibilities

  • Design and own end-to-end LLM evaluation frameworks — including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
  • Build comprehensive test suites for agentic AI systems — validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
  • Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval.
  • Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation — building and maintaining an evolving adversarial test library.
  • Execute fairness, bias, and Responsible AI audits — testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
  • Design and run inference performance benchmarks — measuring latency, throughput, token efficiency, and degradation under peak load — and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
  • Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch).
  • Define the AI testing roadmap and quality standards for the program — establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams.
  • Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one — reviewing prompt architectures, agent designs, and system workflows for testability and risk.
  • Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.

Benefits

  • insurance (including medical, prescription drug, dental, vision, disability, life insurance)
  • flexible spending account and health savings account
  • 16 weeks of new parent leave
  • up to 20 days of bereavement leave
  • 80 hours of Paid Sick and Safe Time
  • 25 days of vacation time
  • 5 personal days
  • 10 annual paid U.S. observed holidays
  • 401k with a best-in-class company match
  • deferred compensation for eligible roles
  • fitness reimbursement or on-site fitness facilities
  • eligibility for tuition reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service