QA AI Automation Engineer

Dynasty Financial PartnersSt Petersburg, FL

About The Position

We are building AI systems that advisors and their clients rely on inside a regulated wealth management platform — an internal AI assistant, an external-facing API layer on top of it, and a growing set of agentic workflows. The bar for correctness, grounding, and auditability is higher here than it is for most AI products. This role exists because traditional QA does not cover that surface. We need an engineer who can evaluate non-deterministic systems, build the tooling that catches AI regressions before they reach an advisor, and hold the line on code quality now that a meaningful share of our code is written by autonomous coding agents. This is a hands-on engineering role, not a test-execution role.

Requirements

  • 5+ years in software quality, test automation, or SDET work, with real ownership of automation architecture — not just test authoring.
  • Demonstrated production experience with LLM-powered systems: agent patterns, prompt engineering, tool/function calling, and orchestration.
  • Hands-on experience with major model APIs (OpenAI, Anthropic, Azure OpenAI) and AI-assisted development tooling (Claude Code, Copilot, Codex or equivalent).
  • Strong programming in at least one of Python, TypeScript/JavaScript, Java, Kotlin, or C#, and comfort reading across the others.
  • Experience building and interpreting evaluation criteria for systems without a single correct answer.
  • Deep CI/CD fluency — pipeline integration, gating, monitoring, and logging.
  • API testing depth and experience validating third-party integrations.
  • Clear written communication; ability to explain quality risks to product owners and root causes to engineers.

Nice To Haves

  • Eval and observability tooling: Langfuse, Promptfoo, OpenTelemetry, New Relic, or equivalents.
  • RAG fundamentals — embeddings, chunking strategy, vector search, retrieval evaluation.
  • Workflow orchestration tooling (n8n or similar).
  • Azure cloud services; .NET ecosystem exposure.
  • BDD/Selenium or comparable UI automation at scale.
  • Financial services, wealth management, or another regulated domain — or a clear appetite for what data governance means in one.
  • Experience mentoring or leading distributed QA engineers.

Responsibilities

  • Own and extend our AI evaluation framework for AI chat, including building and maintaining an evaluation suite for our AI assistant and its API surface, covering correctness, grounding/citation fidelity, retrieval quality, refusal and escalation behavior, tone, latency, and cost per interaction.
  • Curate and version golden datasets and adversarial test sets drawn from real advisor workflows, ensuring they remain representative as the product changes.
  • Design LLM-as-judge and rubric-based scoring where deterministic assertions do not apply, and validate the judges themselves against human-labeled sets.
  • Integrate evaluations into CI to gate prompt, model, retrieval, and tool changes by measured regression.
  • Instrument production traces and close the loop from live failures back into the evaluation suite.
  • Design and operate a multi-agent workflow in our CI/CD pipeline that runs on commit to execute the suite, triage failures, write up defects with reproduction detail, propose or apply fixes, and re-run to verify.
  • Define agent roles, hand-off logic, guardrails, and quality checks to determine when an agent's output is trustworthy enough for auto-application versus routing to a human.
  • Make cost-aware decisions on where an LLM is needed and where deterministic automation is the better tool.
  • Report on the loop's real impact, including autonomous resolution rate, false-positive rate, and engineer hours returned.
  • Define what “good” means for agent-authored code and build automated gates to enforce it, including coverage and mutation testing, static analysis, security and dependency scanning, architectural conformance, and review checklists tuned to how agents fail.
  • Identify and build detection for failure modes specific to agent-generated code, such as plausible-but-wrong logic, silent scope creep, duplicated abstractions, and missing edge-case handling.
  • Partner with architects and delivery teams to maintain velocity without sacrificing quality as agent-assisted development scales across the organization.
  • Partner with product, engineering, and AI Labs to define coverage and acceptance criteria for AI features before they are built.
  • Contribute to release-readiness and go/no-go decisions with evidence.
  • Maintain and modernize our existing automation estate (API, UI, integration).
  • Mentor engineers on best practices.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service