Software Engineer, AI Evaluation

NunaSan Francisco, CA
$147,500 - $232,200

About The Position

Chronic disease isn't managed in a clinic. It is managed at home, in relationships, in the everyday. What's on the dinner table, what gets said, and who notices when someone's struggling. For the 130 million Americans managing a chronic condition, the healthcare system has offered the same answer for decades: a 15-minute doctor's visit, a pamphlet, and a portal login they'll never use. At Nuna, we are building an AI health coach that shows up like a person who actually has time: available at 3am, infinitely patient, and never behind a waiting room. We use motivational interviewing to help patients and their families see themselves clearly, design experiments that fit their real lives, and navigate a system that has not historically been on their side. We are building from the ground up around a simple belief: patients don't want to be healthy; they want their lives back. We're not competing with other health apps. We're competing with the moment a person gives up on getting better. If that's a problem you want to work on, we'd like to talk.

Requirements

  • Significant experience building and shipping reliable production systems and tooling
  • Deep understanding of how to evaluate AI systems - LLM-as-judge, red-teaming and adversarial testing, synthetic scenario generation, and multi-turn and agentic evaluation - and a clear sense of how evals themselves fail. You've deployed evals and automated AI tooling in production, not just prototyped them
  • A testing mindset applied to building the measurement system, not running tests against a spec: adversarial instinct, coverage thinking, regression discipline, and documentation others can build on
  • You use AI in your daily work and build tools that make the people around you more effective
  • Enough fluency in statistics and experimental design to partner with a data scientist on calibration and reliability
  • Can design the workflow and build a functional UI for non-engineers like clinicians and labelers
  • A genuine interest in improving healthcare alongside an interdisciplinary team, with the judgment to tell a launch-blocking issue from a nice-to-have

Nice To Haves

  • Experience in healthcare or another regulated, high-trust domain, and familiarity with the regulatory landscape
  • Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar)
  • Red-teaming or AI safety experience - prompt injection, jailbreaks, adversarial and stress testing
  • Experience with automated, eval-driven model or prompt optimization
  • You've built in an early-stage or fast-moving environment

Responsibilities

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service