Senior QA and AI Evals Engineer

Ruya AIMassachusetts (US) - Onsite, MA
Hybrid

About The Position

Ruya AI is a venture-backed defense-tech company building AI-native intelligent software for sovereign institutions and their affiliated organizations operating in mission-critical, high-stakes environments. Headquartered in Massachusetts (US), with offices across the U.S., Europe, and the Middle East, we are expanding our engineering team with people who demonstrate strong ownership, sound judgment, and the ability to turn complex requirements into reliable products. The Role Quant is the team responsible for our multi-modal intelligence agent. Our software is built to surface hidden patterns across disparate data sources, and generate predictive outcomes for the operators who rely on it. As part of this team, you're responsible for the quality system behind our platform, covering both conventional software behavior and non-deterministic AI outcomes. Our platform must hold up across deterministic user flows and probabilistic AI outputs alike, staying reliable for the operators who depend on it. You'll design and build the automated test suites, evaluation pipelines, and release gates behind that reliability.

Requirements

  • Strong Playwright and API/integration testing experience.
  • Experience with pytest and TypeScript test frameworks such as Vitest.
  • Understanding of asynchronous and distributed-system failure modes.
  • Experience evaluating LLM or other non-deterministic systems.
  • Ability to design high-signal test suites rather than maximizing test counts.

Nice To Haves

  • Experience with Langfuse, playwright-bdd, fault injection, contract testing, security testing, durable workflows, or model/prompt deployment gates.

Responsibilities

  • Build E2E coverage for investigations, maps, graphs, authentication, and chat.
  • Develop API, contract, integration, and component tests across TypeScript and Python services.
  • Create Gherkin/BDD scenarios and build Playwright test cases for important UI flows.
  • Build AI evaluation datasets, metrics, experiments, and regression gates using Langfuse.
  • Test streaming disconnects, replay, cancellation, retries, race conditions, and long-running jobs.
  • Validate authorization boundaries and sensitive-data handling.
  • Integrate test suites into GitHub Actions and manage flaky-test diagnosis.
  • Partner with engineers early to create testable interfaces and useful observability.

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Commuter benefits
  • Relocation assistance
  • Paid time off
  • Leave-of-absence program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service