QA Engineer (SDET) — AI, Data & Platform Quality

ExtractableSan Francisco, CA
Remote

About The Position

This role is for a QA Engineer (SDET) focused on AI, Data & Platform Quality at Finalytics.ai. Finalytics.ai is a leading provider of personalization for the financial industry, using data integrations, machine learning, and real-time technology. The QA role at this company is highly programmatic, going beyond UI testing to focus on testing models, LLMs, agents, and data pipelines. The successful candidate will be responsible for building automated tests, evaluations, and data checks to ensure the quality and trustworthiness of AI-driven personalization features. The role is embedded within the engineering team, sharing the same repository and release flow, and reports directly to the CTO. The tech stack includes Python/Django, JavaScript, MySQL, Celery, BigQuery, and AWS.

Requirements

  • 3+ years in QA/SDET or test automation with a code-first approach.
  • Strong Python skills, including writing clean test code and reading application code.
  • Experience with browser automation (Playwright or Selenium).
  • API and contract testing experience.
  • Genuine interest in testing AI, comfortable with non-determinism, evals, and prompts.
  • Strong SQL skills and experience validating pipelines and reconciling data.
  • Experience building automated quality gates into the deploy and release process.

Nice To Haves

  • Testing or evaluating LLM applications (evals, prompt regression, tool-calling agents, or MCP).
  • Data or analytics QA (BigQuery or ETL/rollup validation).
  • Django, MySQL, or Celery experience.
  • Security testing with SAST/DAST tooling.
  • Familiarity with machine learning.
  • Financial industry, personalization, or CMS/marketing-platform experience.
  • Familiarity with AWS.
  • SaaS startup experience on a fast-moving, multi-tenant platform.

Responsibilities

  • Extend the scenario test runner to capture production personalization requests and replay them across environments, asserting on expected algorithms and content selection, and growing it into automated regression across every client.
  • Write automated tests in Python with pytest across unit, integration, HTTP, and end-to-end tiers.
  • Build headless Playwright end-to-end tests to verify personalized content and tracking rendering on client pages.
  • Harden the pre-deploy quality gate and pre-commit checks to automatically block bad changes.
  • Design evaluations for non-deterministic AI features (conversational analytics assistant, AI content builders, generative SEO), measuring correctness, grounding, and regression across prompt and model versions.
  • Test tool-calling and agentic layers, ensuring function-calling loops pick the right tools and guardrails hold on adversarial input.
  • Validate the agent/MCP interface for contract conformance, rate limiting, authorization, and safe failure.
  • Help set standards for shipping AI, catching hallucinations and drift, and benchmarking prompt/model changes.
  • Build automated data-health checks to flag stale rollups, incomplete coverage, and broken aggregations.
  • Validate data pipelines end-to-end (rollups, funnel/rate/financial ingestion, BigQuery) with drift detection.
  • Guard model inputs to ensure accuracy and completeness of signals for ML.
  • Track platform performance (response times, JS load, page speed) and help maintain speed.
  • Set up quality dashboards for uptime, coverage, data-health, and eval scores.
  • Work in the codebase alongside engineers to diagnose issues across development, release, and deployment.
  • Drive the bug lifecycle: reproduce, capture with a failing test, and verify the fix.

Benefits

  • Direct impact on shaping quality across the platform.
  • Automation-first culture.
  • Remote-first, collaborative, low-ego team.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service