Data Scientist

Headwater ScienceResearch Triangle Park, NC
Hybrid

About The Position

Headwater Science (formerly NoviSci) is a data science and methods company specializing in principled, reproducible evidence generation for complex clinical and regulatory challenges. With deep expertise in comparative effectiveness, causal inference, healthcare utilization and expenditure research, and regulatory-grade analytical software, Headwater Science provides the methodological foundation that delivers reproducible analytic pipelines, novel epidemiologic and statistical methods, and regulatory-grade software validated to hold up under the most demanding scrutiny. The company works with life sciences organizations as a long-term scientific partner. Headwater Science is a Highlander Health company. Learn more at headwaterscience.com. In this role, you will support our research projects by helping to transform source data from healthcare databases into analytic-ready data sets, performing statistical analyses, and generating reports of the results. You will work with small teams of epidemiologists and statisticians whose responsibilities span study design and execution. Your work will focus primarily on executing specifications outlined in study protocols/statistical analysis plans (SAP) and building reproducible analytical pipelines that are understandable, well documented, and compliant with quality control standards. Typical projects include studies of natural history, treatment patterns, and comparative effectiveness/safety, often incorporating negative controls to evaluate treatment group comparability.

Requirements

  • Master's degree, or two or more years of relevant work experience, in biostatistics, epidemiology, health economics, bioinformatics, data science, or a related quantitative field.
  • Strong programming skills in R.
  • Experience querying relational databases, whether directly in SQL or through an interface like dbplyr.
  • Experience writing code that others can read, run, and reproduce.
  • Strong communication, presentation, and collaboration skills with cross-functional technical and scientific teams.

Nice To Haves

  • Experience with real-world healthcare data, such as administrative claims, EHR, hospital chargemaster, or registry data.
  • Experience constructing longitudinal data sets for observational research.

Responsibilities

  • Transform source data from healthcare databases into analytic-ready data sets.
  • Perform statistical analyses.
  • Generate reports of the results.
  • Execute specifications outlined in study protocols/statistical analysis plans (SAP).
  • Build reproducible analytical pipelines that are understandable, well documented, and compliant with quality control standards.
  • Build cohorts and run analyses that produce results.
  • Coordinate through GitLab and stay in close contact as code comes together.
  • Draw on domain knowledge of how each database represents clinical events to make appropriate choices when building study variables.
  • Use dbplyr to query source databases from R, generating SQL against tables that are often too large to hold in memory.
  • Apply eligibility criteria from a study protocol and/or SAP to identify the study population.
  • Determine the index date and start of follow-up for each patient.
  • Build the analytic data set — often one row per patient, with columns for baseline characteristics, changes in treatment and health status over time, outcomes of interest, and censoring.
  • Extend the project pipeline to take the analytic-ready data set through to the finished report.
  • Fit the models specified in the study design, often using causalRisk, our internal package for causal inference with observational data.
  • Run diagnostics on the analysis — such as checking propensity score distributions, weight behavior, and covariate balance — and flag problems back to the study design team.
  • Produce the tables and figures that make up the study results — baseline characteristics, effect estimates, and study-specific outputs like risk curves and treatment patterns — and assemble them into client deliverables.
  • Build project repositories as orchestrated pipelines that reliably reproduce the study results each time they are run.
  • Participate in quality control procedures including code review, testing, and double coding of portions of an analysis as an independent check on the primary implementation.
  • Work closely with the study design team throughout the study lifecycle, incorporating changing requirements, sharing interim results, and providing feedback on how well the study design suits the available data.

Benefits

  • Comprehensive health, dental, and vision coverage for you and your family.
  • 401(k) with company match.
  • Generous PTO and company holidays.
  • Paid parental leave.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service