Bioinformatician I

Duke CareersDurham, NC
$82,159 - $132,722

About The Position

As the Bioinformatician I, you'll play a key role in advancing innovative translational medicine research by helping build and maintain the sophisticated data infrastructure that powers cutting-edge clinical and scientific discovery. Working alongside interdisciplinary teams of researchers, clinicians, engineers, and data scientists, you'll contribute to multi-modal data fusion initiatives that integrate electronic health records, physiological waveform data, and multi-omic datasets. Your expertise will help transform complex healthcare data into meaningful resources that drive impactful research and improve patient outcomes. Join a nationally recognized Department of Surgery and take on some of the most complex data challenges in healthcare research. In this role, you'll help architect scalable systems and high-throughput pipelines that support advanced machine learning, reinforcement learning, and foundation model research across large-scale clinical environments.

Requirements

  • Strong programming expertise in Python and advanced SQL, including large-scale query optimization.
  • Experience designing, developing, and maintaining robust data pipelines in cloud computing environments.
  • Knowledge of distributed data processing frameworks and large-scale data engineering principles.
  • Experience working with electronic health records and complex healthcare datasets.
  • Familiarity with containerization technologies such as Docker and deployment of scalable applications.
  • Strong communication skills with the ability to collaborate effectively with both technical and non-technical stakeholders.

Nice To Haves

  • Experience with OMOP Common Data Model or other healthcare interoperability standards.
  • Experience processing high-frequency physiological waveform data.
  • Familiarity with Azure, AWS, or Google Cloud Platform environments.
  • Knowledge of machine learning infrastructure, reinforcement learning, or foundation model research environments.
  • Experience working with large-scale clinical research initiatives or multi-center collaborations.

Responsibilities

  • Design, build, and optimize distributed data processing workflows utilizing technologies such as PySpark, Dask, and cloud-based computing platforms.
  • Develop and maintain automated, version-controlled ETL/ELT pipelines that support large-scale clinical and research datasets.
  • Process and manage terabyte-scale inpatient healthcare data, including electronic health records, physiological waveform data, and multi-omic datasets.
  • Support multi-site initiatives and federally sponsored programs through rigorous data validation, verification, and quality assurance activities.
  • Contribute to research efforts focused on reinforcement learning methodologies and multi-modal time-series foundation models.
  • Standardize and harmonize complex healthcare data using clinical interoperability frameworks such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model.
  • Collaborate with technical and clinical teams to translate complex data requirements into practical research solutions.
  • Implement scalable cloud-based data management solutions within Azure, AWS, or comparable environments.
  • Utilize optimized data storage formats, including Parquet and HDF5, to support efficient retrieval and analysis of large datasets.
  • Partner with multidisciplinary stakeholders to advance innovative research initiatives across multiple institutions and research networks.

Benefits

  • medical and dental care programs
  • generous retirement benefits
  • family-friendly and cultural programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service