AI Engineer, Decision Intelligence

recoursehealth.comSan Francisco, CA

About The Position

Every dispute is one move in a repeated game against an adaptive opponent. Build the system that gets better at playing it. A federal arbitration system called Independent Dispute Resolution, or IDR, now determines billions of dollars in healthcare payments each year. Providers win the vast majority of disputes, yet most eligible claims are never filed. The process is manual, fragmented, and resource-intensive, and most providers don't have the infrastructure to pursue what they're owed. The No Surprises Act created the framework, and the market already exists. Today it runs on spreadsheets, consultants, and static playbooks. We're building the first intelligent system designed to operate inside it. Automating the paperwork is table stakes. The interesting part is the second half: IDR is baseball-style arbitration, where each side submits one number and an arbitrator picks one. No splitting the difference. That means every submission is a bet, and every outcome is a signal about how to bet better next time. This role owns that.

Requirements

  • You've shipped LLM systems that people depend on. Not demos. Production systems with evals, guardrails, and a real answer for what happens when the model is wrong.
  • You think about systems, not prompts. Prompt engineering is a component. The interesting work is architecture: what's deterministic, what's learned, how they interact, and how the whole thing improves over time.
  • You're rigorous about evaluation. You know that "it seems better" isn't evidence. You build the measurement before you build the feature.
  • You are AI-pilled and current. The landscape moves monthly. You track it because you want to see the next shift before anyone else, and you have opinions about what's real versus hype.
  • You have a bias to action. You don't default to no. Speed of iteration over polish of iteration. You start, you learn, you fix things in motion. Most decisions are reversible and do not need extensive study.
  • You are intellectually honest. You seek out evidence that disconfirms your approach. You say so when you're wrong. You use plain language. You respectfully challenge decisions you disagree with, and once a decision is made, you commit.
  • You put the team first. You are reliable and fully invested. You take your vacations. You check on your teammates. You help build a culture where people do their best work because they are supported, not squeezed.
  • 5+ years engineering, with meaningful recent time building LLM-powered or ML-driven products in production
  • Real experience with evaluation: building golden datasets, running offline evals, measuring whether changes helped
  • Comfort across the stack. You can ship the thing end to end, not just the model layer
  • Strong instincts for where probabilistic reasoning helps and where it introduces unacceptable risk
  • Comfort with TypeScript or Python. The specific stack matters less than the ability to pick things up
  • Sound judgment, technical depth, and ownership mindset are required. Grit matters more than pedigree.

Nice To Haves

  • Experience building in regulated or security-sensitive environments (HIPAA, SOC 2, PCI, financial controls) is a plus
  • A preference for small teams and early-stage chaos over mature org charts
  • Strong plus, not required: healthcare data, claims, or any adversarial decision domain (fraud, risk, pricing, trading). If you have it, you'll move faster. If you don't, we'll teach you.

Responsibilities

  • Build the decision engine that determines which claims to contest and at what offer amount
  • Design the LLM systems that generate arguments: medical necessity, patient acuity, market comparables, the full submission
  • Build evaluation infrastructure. Golden datasets, offline evals, and the ability to know whether a change actually improved outcomes
  • Architect the split between deterministic rules and model reasoning, and defend where you drew the line
  • Close the feedback loop from arbitration outcomes back into the system, so every decision makes the next one better
  • Build for explainability. Every recommendation needs reasoning a customer would accept
  • Partner with operations and payer strategy, who see patterns in the claims before the data does

Benefits

  • Institutional backing
  • Shared platform team spanning engineering, strategy, design, and back-office
  • Early access to large provider systems
  • Funded, validated opportunity with real customers and real data
  • Small, nimble team
  • Value clarity over theater
  • Opportunity to have the best work of your career
  • Supportive culture
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service