Founding Research Engineer

Nolla Health•New York, NY
•$180,000 - $240,000•Onsite

About The Position

Nolla is building AI doctors to make top-quality healthcare accessible to everyone. Nolla Skin is a leading medical skincare treatment app, and they have expanded into urgent care with NollaMD and are developing specialty apps. The company has raised $4.5M in pre-seed funding from General Catalyst. This role involves working on evaluation infrastructure and data loops for AI clinical models, directly contributing to patient care and research publications. The engineer will collaborate with the founding team and clinical partners, writing code daily, designing studies, and publishing research. The work includes building and extending benchmarks for clinical accuracy, safety, and documentation, developing grading systems, and creating data loops with clinicians for case conversion and review. The role also involves leading clinical evaluations from protocol to publication and collaborating on reward design and data curation.

Requirements

  • Built eval infrastructure for LLM systems that others depended on.
  • Understanding of contamination, difficulty calibration, judge bias, and reward hacking.
  • Ability to write a paper, ideally with authorship on empirical ML work, human comparison, or benchmark release.
  • Excitement to work with physicians and translate their judgment into model performance.
  • Experience working on small teams and building things from zero.

Nice To Haves

  • Clinical AI evaluation experience: rubric-based health benchmarks, simulated-patient studies, agentic clinical benchmarks.
  • Post-training experience (RLHF, DPO, GRPO or similar).
  • Experience with multimodal models.
  • Familiarity with eval and environment frameworks such as Inspect, Verifiers, or Harbor.

Responsibilities

  • Design and maintain benchmarks for judging clinical accuracy, safety, and documentation quality.
  • Extend benchmarks to agentic systems, including tool use, long-horizon tasks, patient history memory, uploaded records, labs, and escalation behavior.
  • Build the grading stack, including deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings.
  • Improve the production harness used by the product.
  • Turn real, consented cases into eval cases and training examples with clinician review, de-identification, and provenance.
  • Run clinician review workflows to produce rubrics, labels, and feedback at scale.
  • Lead clinical evaluations from protocol through publication and publish benchmarks.
  • Collaborate with post-training teams on reward design and data curation, and run experiments.

Benefits

  • Medical, dental, and vision insurance
  • HSA/FSA eligible
  • Flexible, unlimited vacation
  • Meal stipends
  • Team retreats
  • Meaningful equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service