Machine Learning Data Engineer (AIM2)

PfizerCambridge, MA
$106,000 - $176,600Hybrid

About The Position

We are seeking an experienced Machine Learning Data Engineer to develop, operationalize, and continuously improve production bioinformatics pipelines and data products that support AI-driven discovery within the Internal Medicine Research Unit (IMRU). The successful candidate will integrate and curate complex biological and clinical datasets for machine learning model development and deployment, while building scalable cloud-based data pipelines that enable insight generation across IMRU's proprietary and external data resources. This role will partner closely with the Integrative Biology organization and the AI for Internal Medicine (AIM2) Discovery Group to deliver scientifically robust, trusted, and reusable data capabilities that accelerate evidence generation and improve the speed and quality of research across the portfolio.

Requirements

  • PhD in Computational Biology, Biology, Physics, Statistics, or a related technical discipline
  • Masters in Computational Biology, Biology, Physics, Statistics, or a related discipline and a minimum of two years of experience developing data products and data integration solutions in a research or industry environment
  • Single‑cell/NGS, functional genomics, genetics, or proteomics data analysis experience
  • Hands‑on experience developing Nextflow pipelines for processing NGS data
  • Strong full‑stack programming skills with a focus on Python
  • Experience solving complex analyses/problems in a timely fashion
  • Excellent communication and collaboration skills with experience working effectively in cross-functional teams
  • Permanent work authorization in the United States.

Nice To Haves

  • Background or demonstrated interest in life sciences, pharmaceutical research, drug discovery, or bioinformatics.
  • Proven expertise in software engineering best practices, including python package development, DevOps, cloud architectures, CI/CD, and engineering tooling
  • Hands-on experience handling, processing, integrating, and analyzing large heterogenous data sets data in a drug discovery research environment
  • Experience with Claude Code or equivalent
  • Strong publication record with demonstrated contributions to the field

Responsibilities

  • Developing, deploying, and operating production‑grade Nextflow pipelines on cloud infrastructure, ensuring scalable and reproducible execution.
  • Owning pipeline lifecycle management, including upgrades, troubleshooting, performance tuning, and reliability improvements for reusable workflows.
  • Implementing DevOps best practices for pipelines and platform services (e.g., CI/CD, automation, and engineering tooling).
  • Partnering with wet‑lab and research scientists to translate data analysis requirements into robust, production‑ready pipeline and platform solutions.
  • Processing, organizing, harmonizing and integrating relevant data types (e.g.single-cell omics, bulk transcriptomics, ATACseq, proteomics, functional genetics, and genetics) as inputs for ML model training or inference.
  • Developing and evolving an omics data platform that enables efficient, scalable processing and delivery of omics datasets as reliable data products.
  • Driving collaborations with external partners and vendors to strengthen pipeline quality, sustainability, and adoption of best practices.

Benefits

  • 401(k) plan with Pfizer Matching Contributions
  • Additional Pfizer Retirement Savings Contribution
  • Paid vacation, holiday and personal days
  • Paid caregiver/parental and medical leave
  • Health benefits to include medical, prescription drug, dental and vision coverage
  • Relocation assistance may be available
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service