Senior Software Test Engineer

Curative HR LLCAustin, TX

About The Position

We are hiring a Senior Software Test Engineer to help us deliver quickly with the right level of quality. If you're picturing test plans, sign-off gates, and sprint ceremonies, this is not that job. At Curative, every engineer owns what they ship. Your job is to make that possible at speed. A big part of that is our AI agents, which do real operational work in production. We're investing in the evaluation layer to match: eval suites, regression harnesses, and systematic measurement. But the role is not only infrastructure. You'll work directly with business teams, become a domain expert in corners of our business, test the way real users work, and catch the gap between what was asked for and what was shipped. Some days you're writing an eval harness; some days you're hunting a bug an operations teammate can feel but can't pin down; some days you're shaping what gets built next. You'll advise on quality concerns, and you'll jump in and fix things yourself: same day, not next sprint. If ambiguity sounds stressful, this is the wrong role. If moving between code, product, and people in the same afternoon sounds like the best part of the job, keep reading. At Curative, AI writes most of the code. Engineers direct it, using agentic AI coding tools as the primary development surface: setting context, making the decisions the AI cannot, and keeping the bar high on what ships. This is not a role for someone who wants to hand-roll every line, nor for someone who will accept whatever the AI produces. We want the engineer in between: fundamentals strong enough to catch a wrong answer fast, discipline to review every diff, and ambition to drive several times the output of a traditional IC.

Requirements

  • 5+ years in software quality, testing, or engineering with experience in both automation and exploratory testing.
  • Fluency with LLM evaluation techniques, including golden datasets, LLM-as-judge, programmatic graders, and regression suites for prompts and agent loops.
  • Hands-on experience with AI agents and understanding of common failure modes like silent drift, tool misuse, and compounding errors.
  • Demonstrated ability to find obscure bugs and a track record to prove it.
  • Proficiency in debugging using logs, traces, queries, and production forensics.
  • Ability to fix bugs independently in Python or TypeScript.
  • Strong communication skills with business stakeholders, including the ability to run sessions with non-technical leads and extract necessary information.
  • Comfort with ambiguity and a fast-paced environment, able to start tasks without formal tickets or sprint boundaries.
  • Pragmatic quality judgment, able to determine acceptable risk levels for shipping.
  • Sharp written communication skills that influence action.
  • Primary development tool is an AI coding assistant like Claude Code or Cursor.
  • Experience using AI to accelerate quality work (eval cases, graders, triage) and understanding the limits of AI judgment.
  • Discipline to review every AI-generated diff and not merge code based on assumptions.

Nice To Haves

  • Experience building eval or observability infrastructure from scratch that was successfully adopted.
  • Familiarity with LLM observability tooling such as LangSmith, Braintrust, Arize, or homegrown solutions.
  • Experience defining unsupervised agent capabilities and building guardrails.
  • Experience acting in a product role, including defining requirements, behavior, and serving as a de facto Product Manager.
  • Experience building deep domain expertise from scratch in a complex operational business.

Responsibilities

  • Build and maintain the eval platform, including harnesses, datasets, graders, and CI integration to measure agent behavior (task success, tool-use correctness, drift, latency, cost).
  • Perform hands-on, exploratory testing of the product from the perspective of members, operations teams, and providers, including edge cases.
  • Validate that the correct product is built by working with business stakeholders to understand workflows and identify discrepancies between requirements and shipped features.
  • Develop domain expertise in specific areas of the business to be a valuable resource for teams.
  • Analyze production signals, identify failure taxonomies, and create dashboards to diagnose issues, reproduce failures, find root causes, and fix them or provide precise diagnoses.
  • Exercise product judgment when necessary, including writing requirements, proposing behavior, and making decisions in the absence of a Product Manager.
  • Prioritize efforts in a dynamic environment, determining where to focus first, which agents carry the most risk, and differentiating between critical and minor failures.

Benefits

  • Curative Health Plan (100% employer-covered medical premiums for employee and 50% for dependents on the base plan)
  • $0 copays and $0 deductibles (with completion of Baseline Visit)
  • Preventive and primary care
  • Mental health support
  • One-on-one care navigation
  • Chronic condition programs (diabetes, weight, hypertension)
  • Maternity and family planning support
  • 24/7/365 Curative Telehealth
  • Pharmacy benefits
  • Comprehensive dental and vision coverage
  • Employer-provided life and disability coverage with supplemental options
  • Flexible spending accounts
  • Generous PTO policy
  • 11 paid annual company holidays
  • 401K for full-time employees
  • Generous 8–12 weeks paid parental leave, based on role eligibility.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service