Data Scientist Pulmonary Research

Mass General BrighamBoston, MA
Hybrid

About The Position

The position will work directly in collaboration with Dr. Shappell, a critical care physician researcher with expertise using data from Electronic Health Records (EHRs) to study the epidemiology and outcomes of critical illness. The position will primarily be responsible for the creation, maintenance, and ongoing operations of the new MGB Critical Care Data Repository, which aims to provide highly granular research-ready ICU data to MGB investigators for clinical operations, quality improvement, and research projects. This will include working with existing MGB analysts to complete set up of EHR data pipelines and databases using Clarity and MGB’s cloud-computing environment, Snowflake; collaborating with clinical experts on data mapping and cleaning; performing in-depth data quality assurance including the production of visualizations and reports; and providing requested data to PCCM Division leadership as needed. In addition, the Data Scientist will support Dr. Shappell’s health services research on critical illness including sepsis, shock, severe viral infections, and respiratory failure; they will also support the full integration of the MGB data repository into an established critical care research consortium via application of a novel common data model. This work will involve completing and ensuring compliance with latest version of common data model, developing and coding new research projects that utilize consortium data, running code locally from other consortium PI-led projects, and attending weekly consortium meetings. There is opportunity for expansion into work with other EHR data modalities, including unstructured data from clinical notes, waveform data, and/or imaging data, if desired. Responsible for analyzing data, uncovering the underlying data patterns and logic, and developing data-driven applications. They will work towards developing solutions for the entire problem-solving cycle: find/prioritize the problems, research the best algorithms to solve the problems, design robust, practical solvers, and implement them.

Requirements

  • Bachelor's Degree Related Field of Study required or Master's Degree Related Field of Study preferred
  • Experience working in data science-type positions and with large data sets
  • Technical expertise with SQL, Stata, R, Python.
  • Experience with database creation and engineering.
  • Ability to create reports and dashboards with Tableau.
  • Skilled in data analysis with Python and SQL.
  • Knowledge of statistics.
  • Knowledge of machine learning.
  • Practiced knowledge of numerical optimization algorithms.

Nice To Haves

  • Preferred SQL database management and administrative knowledge.

Responsibilities

  • Analyze complex and high-dimensional clinical and operational data for dependencies, patterns, outliers, inaccuracies, and validity.
  • Apply knowledge of statistics, machine learning, programming, data modeling, simulation, and/or advanced mathematics to recognize patterns, identify opportunities, pose business questions, and make valuable discoveries.
  • Contribute to the design and evaluation of optimal data-driven metrics (such as physician/facility performance criteria, bottleneck metrics, productivity limits, processing delay/error reports, etc.).
  • Use a flexible, analytical approach to design, develop, and evaluate predictive models and advanced algorithms that lead to optimal value extraction from the data.
  • Generate and test hypotheses and analyze and interpret the results.
  • Leads creation and maintenance of databases in Clarity’s cloud-computing environment, Snowflake, using SQL.
  • Data cleaning and management.
  • Prepares analytic datasets using data analysis software (E.g. R tidyverse packages, Pandas for Python) from electronic health care record data sets.
  • Follows the standards of transparent and reproducible data science including use version control software such as git and code with a literate programming style.
  • Anticipates data management needs, including identifying when data pulls will be necessary and obtaining newer versions of datasets.
  • Data visualization with appropriate statistical software, such as the ggplot2 package in R.
  • Reports results of data analysis, prepares figures and tables.
  • Supports PI and trainees in the drafting and revision of manuscripts and grant proposals.
  • Collaborates effectively with local and external PIs and data analysts/scientists.
  • Performs other related work as needed.

Benefits

  • comprehensive benefits
  • career advancement opportunities
  • differentials
  • premiums
  • bonuses
  • recognition programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service