Founding Research Engineer

Nolla Health•New York City, NY
•Onsite

About The Position

Nolla is building AI doctors to make top-quality healthcare accessible to everyone. Nolla Derm is a leading medical skincare app, and they have expanded to NollaMD for urgent care and are developing specialty apps. The company has raised $6.5M from General Catalyst and other investors. This role involves working on evaluation infrastructure and data loops for AI clinical models, directly with the founding team and clinicians. The work directly impacts patient care and involves hands-on coding, study design, and publication.

Requirements

  • Experience building eval infrastructure for LLM systems that others depended on.
  • Understanding of contamination, difficulty calibration, judge bias, and reward hacking.
  • Ability to write a paper, with authorship on empirical ML work, ideally with human comparison or benchmark release.
  • Excitement to work with physicians and translate their judgment into model performance.
  • Experience working on small teams and building things from zero.

Nice To Haves

  • Clinical AI evaluation experience, including rubric-based health benchmarks, simulated-patient studies, and agentic clinical benchmarks.
  • Post-training experience (RLHF, DPO, GRPO or similar).
  • Experience with multimodal models.
  • Familiarity with eval and environment frameworks such as Inspect, Verifiers, or Harbor.

Responsibilities

  • Design and maintain benchmarks for clinical accuracy, safety, and documentation quality.
  • Extend benchmarks to agentic systems, including tool use, long-horizon tasks, patient history memory, and escalation behavior.
  • Build the grading stack, including deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings.
  • Improve the production harness used by the product.
  • Develop the data loop with clinicians to turn consented cases into evaluation cases and training examples, including review, de-identification, and provenance.
  • Run clinician review workflows for rubrics, labels, and feedback at scale.
  • Lead clinical evaluations from protocol through publication.
  • Publish benchmarks for the field.
  • Collaborate with post-training teams on reward design and data curation.
  • Run experiments to improve model performance.

Benefits

  • Meaningful equity, commensurate with experience
  • Medical, dental, and vision insurance
  • HSA/FSA eligibility
  • Flexible, unlimited vacation
  • Meal stipends
  • Team retreats
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service