Staff Applied Scientist, Reinforcement Learning

Hippocratic AIMenlo Park, CA
Onsite

About The Position

LLM post-training is where raw capability becomes reliable, safe behavior — and in healthcare, the stakes are as high as they get. You'll own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline end to end, to improve our models' clinical reasoning, safety, and alignment. Your models will be deployed to interact with millions of patients across diverse clinical use cases.

Requirements

  • MS or PhD in CS or relevant field
  • 5+ years or experience in NLP, LLM training, or RL
  • 2+ years experience in RL for LLM post-training
  • Experience with large-scale (50B+ parameter and multi-node) LLM training
  • Strong Python and PyTorch coding skills
  • Experience with RLHF, RLVR, LLM-as-judge or similar methods for LLM post-training

Nice To Haves

  • Publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
  • Healthcare domain experience

Responsibilities

  • Design RL and OPD post-training methods (RLHF, RLVR, OPD, etc.)
  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines
  • Develop conversational AI environments and simulations for healthcare RL training with synthetic data
  • Automate post-training loops with agents (auto-research)
  • Run rigorous experiments to understand what drives post-training gains
  • Collaborate with research, engineering, and clinical teams
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service