NLP / LLM Data Scientist

Dandelion Health Inc
$135,000 - $165,000Remote

About The Position

Dandelion Health is building the world’s largest AI training and clinical development platform. We make data access easy for AI developers, pharma, and medical devices, while raising the bar for patient safety and data quality. Our mission is to be the go-to platform for healthcare organizations to build responsible clinical AI products. We pride ourselves on learning from data and improving to help our clients enhance health through AI. We partner with health systems to safely and ethically make their de-identified patient data available to AI developers, currently acquiring data from Sharp HealthCare, Sanford Health, and Texas Health Resources, with more joining soon. Our data includes clinical information dating back to July 1, 2016, representing over 10 million patients, encompassing structured data (EMR, claims), unstructured text (clinical notes, radiology reports), images (DICOM, pathology), video, waveforms, and continuous streaming monitoring data. As a healthcare data scientist, you will leverage and build upon large language models (LLM) and other ML-based approaches for meaningful data abstraction from unstructured and structured healthcare data. You will join a team of data scientists responsible for creating and maintaining AI-ready datasets. Your work will involve curating datasets by identifying patient subpopulations or disease cohorts, developing methodologies to abstract information from multimodal healthcare data sources for patient phenotyping, and pooling data for rapid exploratory AI/ML analyses, model experimentation, and validation. You will utilize your data expertise, programming abilities, and critical thinking skills to support the technical product team and develop your own analyses to derive insights and enhance datasets based on use cases. The ultimate goal is to deliver high-quality data to clients building products that improve patient health. You will report to the Data Science Manager, under the Head of Data.

Requirements

  • Advanced degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics), or B.S. with at least 5 years of professional experience
  • At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data by applying advanced machine learning and AI techniques.
  • Prior experience with clinical and healthcare data is a strong bonus.
  • Fluency in Python and SQL, including fluency with ML/NLP libraries (PyTorch, Tensorflow, HuggingFace, etc.)
  • Familiarity with using modern applied LLM techniques on real-world data
  • Strong technical writing, editing, and communication skills, along with a collaborative mindset
  • Excellent organizational skills with an ability to embrace change and effectively manage multiple projects and consistently plan work to meet deadlines

Nice To Haves

  • Experience working in or with startups is a plus
  • Git and version control
  • Familiarity with encryption methods
  • Prior experience querying EDWs or databases and creating reports or analytics for healthcare data
  • Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts
  • Any medical ontology experience
  • Any experience working with DICOM or other imaging modalities
  • Experience with AWS
  • Experience with publishing work in peer-reviewed journals

Responsibilities

  • Develop Natural Language Processing (NLP), Large Language Model (LLM) and other ML-based pipelines to abstract relevant labels from text-based healthcare data and store them in scalable data models;
  • Query complex source systems in a range of health data sources (e.g., EMRs, semi-structured reports, free-text clinical provider notes) to identify key data elements and create and enrich high-quality datasets for real-world evidence analyses and training AI algorithms;
  • Own data extraction, wrangling, labeling and QC tasks to create analytical datasets that include abstracted clinical concepts and provide a range of solutions to support customers’ AI activities;
  • Stay current on the latest in applied NLP and generative AI methods and proactively leverage these technologies where applicable;
  • Support the design, testing, validation, analysis, and merging of multimodal data structures from a wide variety of source systems;
  • Develop code and documentation to deliver high-quality and HIPAA-compliant data products on time to customers;
  • Identify and resolve problems using your knowledge, background, and troubleshooting skills;
  • Ensure accuracy, data integrity, and validity of data and analysis in all work;
  • Provide support for technical product team to advance development of the suite of data-related product offerings;
  • Summarize the complexity of abstraction methods, findings and recommendations into clear explanations and presentations for internal and external audiences that have a varying range of technical and clinical experience;

Benefits

  • Remote work and flexible hours.
  • Complete wellness benefits including healthcare, dental, vision, PTO, sick days and more.
  • Professional development days to build your skills
  • Collegial work environment
  • Academic bent towards inquiry and problem solving but start-up speed and flexibility
  • Great balance of focus time to work on projects but easy to access team members to discuss issues and work collaboratively
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service