Principal Data Engineer, Biologics Discovery

Johnson & Johnson Innovative MedicineHopewell Township, NJ
$117,000 - $201,250Onsite

About The Position

Johnson & Johnson Innovative Medicine is seeking a Principal Data Engineer dedicated to our Biologics Discovery organization. This is a high-leverage role responsible for shaping how discovery data is structured, connected, and made AI-ready across Biologics Discovery. The role serves as the bridge between scientific workflows, data consumers, and technology partners, ensuring that discovery data products support scientific research, analytics, machine learning, and agentic workflows. This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ. (No remote option.) Why this role matters: High-quality, well-governed scientific data is foundational to our vision for AI-enabled biologics discovery. This role provides senior technical leadership within Biologics Discovery, translating scientific needs into data products, scientific data models, and requirements, and working with enterprise data and technology partners to ensure discovery data is trusted, connected, and AI-ready as the broader ecosystem evolves. Position Summary As a Principal Data Engineer, you will lead the design of discovery data products, scientific data models (schemas, entities, and relationships), and integration requirements that enable discovery data to be acquired, connected, harmonized, and delivered across the Biologics Discovery ecosystem. You will work closely with scientists and AI/ML teams to translate their needs into durable, reusable, and AI-ready data assets. Working in close partnership with enterprise Data Strategy & Products and Technology teams, you will ensure discovery data needs are represented in enterprise standards and that those standards are effectively applied within Biologics Discovery. You will help shape the future-state discovery data ecosystem while delivering near-term value through trusted data products, harmonized data, and metadata practices that support long-term interoperability and reuse.

Requirements

  • Degree in Computer Science, Data Science, Engineering, or a related computational field.
  • 8+ years (Bachelor's), 5+ years (Master's), or 3+ years (Ph.D.) of experience designing and delivering data products, data models, and analytics-ready datasets within pharmaceutical, biotechnology, or life sciences organizations.
  • Deep proficiency in Python and SQL, with hands-on experience designing and implementing reusable and scalable data assets on cloud data platforms (e.g., Snowflake, AWS, Azure, BigQuery) to support analytics, machine learning, and AI-driven workflows.
  • Experience applying FAIR data principles, metadata management, controlled vocabularies, data lineage, and provenance practices in scientific data environments.
  • Demonstrated technical leadership through architecture reviews, mentorship, code reviews, or leadership of complex technical initiatives.
  • Proven ability to lead complex cross-functional technical initiatives and establish data standards across multiple stakeholder groups.
  • Strong communication and the seniority to represent the team credibly in architecture, data, and strategy forums.

Nice To Haves

  • Prior technical mentorship or leadership responsibilities.
  • Experience with ontologies, semantic technologies, knowledge graphs, or ontology-driven data architectures.
  • Experience in biologics discovery, high-throughput experimentation, or external partner (CRO/CDMO) data integration.
  • Experience with modern software and data engineering practices, including automated testing, CI/CD, containers, orchestration, and monitoring.

Responsibilities

  • Design and deliver AI-ready discovery data products that support ML, AI, and insight generation across Biologics Discovery, applying agile delivery practices to respond to evolving scientific needs.
  • Define and lead the delivery of scalable integration requirements, transformation patterns, and data schemas that support discovery data acquisition, harmonization, and downstream analytics, working with scientific, data, and technology stakeholders to enable reliable data exchange across systems.
  • Translate scientific and analytical requirements from discovery teams into data product specifications, data contracts, acceptance criteria, and delivery requirements, in partnership with scientists, AI/ML teams, and technology partners.
  • Define access and data consumption patterns that enable analytics, modeling, and agentic AI workflows, aligned with industry data standards and frameworks.
  • Catalog discovery instruments, data types, and data sources to inform data product prioritization and sustainable integration approaches.
  • Establish, champion, and drive adoption of standards and best practices for discovery data products, including data quality, provenance, lineage, reproducibility, metadata, and documentation.
  • Partner with ontology, data architecture, platform, and AI teams to ensure discovery data products are connected, discoverable, and suitable for advanced analytics, ML, and agentic AI applications.
  • Apply FAIR data principles, so data products are reusable, scalable, and interoperable.
  • Serve as a thought leader in scientific data architecture, harmonization, and AI-ready data practices across the organization.

Benefits

  • medical, dental, vision, life insurance, short- and long-term disability, business accident insurance, and group legal insurance.
  • consolidated retirement plan (pension) and savings plan (401(k)).
  • Vacation – up to 120 hours per calendar year
  • Sick time - up to 40 hours per calendar year
  • Holiday pay, including Floating Holidays – up to 13 days per calendar year
  • Work, Personal and Family Time - up to 40 hours per calendar year
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service