Senior / Principal Data Infrastructure Engineer

Flagship Pioneering•Cambridge, MA

About The Position

We are seeking a highly skilled, hands-on Senior / Principal Data Infrastructure Engineer to design, build, and own the data systems that make our proteomics data discoverable, trustworthy, and usable across the organization. Expedition Medicines develops machine learning tools trained on proteomics data generated by our experimental platform, and this role is critical to the robust, scalable collection, management, and analysis of that data at scale. This is a senior individual-contributor role for an engineer who writes the code, ships the systems, and sets the technical direction. You will work independently, with a clear point of view on where our data infrastructure needs to go, and you will turn that vision into working, production-grade systems that ML scientists, medicinal chemists, and biology teams rely on every day.

Requirements

  • Bachelor's degree in Computer Science, Data Engineering, Computational Biology, Proteomics, or a related technical field; advanced degree is a plus.
  • 8+ years of hands-on experience in data infrastructure, data engineering, informatics, or related fields, preferably in proteomics, bioinformatics, or other biotech environments.
  • Proven track record of independently designing, building, and operating terabyte-scale data systems in a biotech or life sciences environment.
  • Demonstrated ability to set technical direction and deliver complex projects with minimal oversight.
  • Expertise in cloud infrastructure (e.g., AWS, Azure), Kubernetes, and infrastructure management as code tools (CFN, CDK, Terraform, ARM).
  • Deep understanding and experience of data modeling with database and data warehouse systems (e.g., Postgres, Redshift, Snowflake).
  • Strong familiarity with lakehouse and medallion-style architectures (e.g., Delta Lake, Iceberg).
  • Experience with data pipeline architecture and workflow orchestration (e.g., Flyte, Dagster, Airflow), including APIs, schedulers, and software integration.
  • Strong proficiency in Python, SQL; comfortable with R.
  • Hands-on experience with DevOps, software design lifecycle (SDLC), and automation tools (e.g., Terraform) and CI/CD practices.
  • Strong written and verbal communication skills, with the ability to explain technical decisions to both technical and non-technical colleagues.

Nice To Haves

  • Experience using LLMs and agentic tooling (e.g., Claude, Model Context Protocol servers) to build interactive, natural-language query interfaces that give scientific teams direct access to data is highly preferred.
  • Experience with mass spectrometry and/or NGS data sets and biomarker discovery workflows.
  • Experience with LIMS and data management systems (e.g., Dotmatics, CORE LIMS).
  • Familiarity with laboratory automation systems, including integration with robotic and sample management platforms.

Responsibilities

  • Architect and build: Design and implement the core data platform that integrates experimental proteomics data with computational tools, with high availability, scalability, and security.
  • Write production code.
  • Own the technical vision: Define and drive the technical roadmap for data infrastructure, identifying the highest-leverage problems.
  • Automate and integrate: Build automated workflows across proteomics research environments, including high-throughput assays and mass spectrometry data processing.
  • Integrate instruments, pipelines, and LIMS (e.g., Dotmatics) for seamless, traceable data capture.
  • Data pipelines: Develop and operate robust data pipelines using workflow orchestration tools (e.g., Flyte), and design layered (e.g., medallion-style) data architectures turning raw instrument output into ML-ready datasets reliably and reproducibly.
  • Cloud infrastructure: Design, deploy, and maintain cloud-based infrastructure for biological and proteomics data processing, storage, and analysis, using infrastructure-as-code and CI/CD to enable continuous improvement of data systems.
  • Data integrity and quality: Establish and implement rigorous standards for data integrity, lineage, traceability, and consistency.
  • Build the tooling that enforces best practices for data capture, storage, and sharing across manual and automated workflows.
  • Partner with scientists: Work directly with our proteomics, chemistry, biology and machine learning teams to translate scientific objectives into data infrastructure solutions.
  • Be the go-to technical expert on data systems.
  • Engineering practices: Set engineering standards through code review, design review, and example.
  • Mentor engineers and scientists informally and share knowledge.

Benefits

  • healthcare coverage
  • annual incentive program
  • retirement benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service