About The Position

The Senior QA Engineer will be responsible for assuring the quality of our products, creating automated tests, and managing builds and continuous integration. You will automate tests for microservices and their related front ends, and manage infrastructure across development and production cloud environments. This role goes beyond conventional application testing. Our platform combines graph data stores, graph analytics, machine learning models, and agentic AI workflows — systems whose outputs are probabilistic, whose correctness is contextual, and whose decisions carry regulatory weight. You will help define what “passing” means for components that don’t return the same answer twice, and build the harnesses that hold them to it.

Requirements

  • A Bachelor’s Degree in Computer Science or a related field.
  • Proven track record of extensive hands-on experience in automated testing of web applications and services, including designing, developing, and executing automated tests in Python using a modern browser automation framework (Playwright, Selenium WebDriver, or equivalent).
  • Demonstrated understanding of maintainable test architecture — Page Object Model or equivalent patterns — for end-to-end testing
  • Extensive hands-on experience on Python scripting
  • Experience testing web services and REST APIs (Postman, Swagger)
  • Working knowledge of graph databases — Neo4j, Amazon Neptune, TigerGraph, JanusGraph or similar — including writing and validating graph queries
  • Exposure to testing AI/ML-backed features, and a clear grasp of why probabilistic systems need different assertions than deterministic ones
  • Scripting experience beyond Python: Unix shell, Ruby, or similar
  • Continuous integration and pipeline-based delivery (Jenkins, GitHub Actions, GitLab CI, or similar)
  • Experience with SQL
  • Experience recommending process improvements

Nice To Haves

  • Hands-on experience with agentic frameworks — LangChain, LangGraph, Model Context Protocol (MCP), or comparable orchestration tooling
  • Familiarity with LLM evaluation tooling and practices: golden datasets, rubric-based scoring, hallucination and jailbreak testing, red-teaming
  • Graph analytics libraries such as Neo4j GDS, Apache Spark GraphX, or NetworkX
  • Observability-driven testing — synthetic monitoring, canary analysis, and validation through traces, logs and metrics
  • Docker, Kubernetes, Terraform, Ansible; microservices, container deployment and service orchestration
  • Cloud APIs (AWS, Azure, GCP)
  • JIRA for issue tracking and escalation

Responsibilities

  • Build and maintain end-to-end and component test suites using modern browser automation with auto-waiting, trace-based debugging, and parallel execution as the default
  • Shape coverage around the test pyramid rather than the UI — API and contract tests as the primary safety net, with UI reserved for genuine user journeys
  • Embed quality gates directly in CI/CD pipelines — sharded parallel runs, risk-based test selection, and pass/fail criteria that block promotion without manual sign-off
  • Provision ephemeral, containerised test environments on demand through infrastructure-as-code, with synthetic and masked test data generated per run instead of maintained by hand
  • Treat test reliability as a product: track flake rates, quarantine and fix unstable tests, and hold suites to explicit stability and runtime SLAs
  • Use AI-assisted tooling for test authoring, coverage gap analysis, and locator resilience — while owning the judgement call on what the generated tests are actually worth
  • Work shift-left alongside engineers — contributing test coverage with the feature, not after it — and partner on triage to drive issues to root cause
  • Design and automate validation for graph database layers — schema integrity, relationship correctness, and query behaviour across Cypher, Gremlin, or SPARQL
  • Build regression coverage for graph analytics outputs: pathfinding, centrality, community detection, and link analysis used in fraud-ring identification and network risk scoring
  • Validate graph ingestion and transformation pipelines for data completeness, deduplication, and referential accuracy at scale
  • Develop evaluation harnesses for ML model behaviour — accuracy thresholds, drift detection, bias and fairness checks, and explainability outputs required for regulatory review
  • Build automated test suites for agentic workflows: tool-calling correctness, multi-step orchestration, context handling, failure recovery, and guardrail enforcement
  • Design assertion strategies for non-deterministic outputs, including golden-set comparison, semantic similarity scoring, and LLM-as-judge evaluation
  • Establish prompt and model regression testing so behavioural changes are caught before release
  • Validate feature pipelines and training/serving consistency in partnership with data science teams
  • Incorporate best practices into dev/test/deploy processes and recommend improvements
  • Support application onboarding to the defined target operating model to achieve build and release automation
  • Document, track, and escalate issues as appropriate

Benefits

  • comprehensive health and wellness plans
  • paid time off
  • company holidays
  • flexible and remote-friendly opportunities
  • maternity/paternity leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service