Principal Engineer 1, Data Platforms

HalozymeSan Diego, CA

About The Position

The Principal Engineer 1, Data Platforms serves as the technical lead accountable for the data platform: architecting the solution, building alongside the team, holding the quality bar, and driving delivery with urgency and accountability. They work directly with business partners across the organization to understand their data needs, translate requirements into platform capabilities, and ensure the UDP delivers real, measurable value to every function it serves, and build the operational foundation that keeps the platform healthy as adoption scales: supporting models, SLA frameworks, and feedback loops that make the platform responsive and reliable. This role helps define and track technical KPIs like pipeline reliability, data freshness, query performance, onboarding velocity, incident resolution and use them to drive continuous improvement.

Requirements

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or a related technical field with 10+ years experience of progressive experience in data engineering, data platform architecture, or analytics infrastructure roles
  • 2+ years technically leading data engineers or analytics engineers in a platform or lakehouse environment
  • Deep, hands-on expertise with Microsoft Fabric (lakehouses, notebooks, Data Factory, deployment pipelines); equivalent depth in Databricks or Azure Synapse considered with demonstrated willingness to go deep on Fabric
  • Strong production-level proficiency in PySpark, T-SQL, Python, and modern ELT/ETL design patterns
  • Proven experience designing and operating medallion-architecture (bronze/silver/gold) or equivalent layered data platforms at enterprise scale
  • Demonstrated ability to manage external implementation vendors at a technical level — reviewing code, challenging architecture decisions, and holding delivery quality, not just tracking timelines
  • Experience with DevOps and CI/CD for data platforms (Azure DevOps, GitHub Actions, Fabric deployment pipelines, infrastructure-as-code)
  • Familiarity with data governance tooling (Microsoft Purview, Unity Catalog, or comparable) and data quality frameworks

Nice To Haves

  • Master’s degree preferred
  • An equivalent combination of experience and education may be considered
  • Biopharmaceutical, life sciences, or regulated industry experience strongly preferred — including familiarity with GxP data handling, validation requirements, and audit readiness
  • Experience integrating pharma-specific source systems such as Master Control, Veeva, LIMS, ELN such as Benchling, CTMS, EDC, or HRIS platforms is a significant plus
  • Exposure to AI/ML workloads on a Lakehouse (feature stores, model training pipelines, vector databases) is a plus
  • Microsoft Fabric Data Engineering Certificate is a plus

Responsibilities

  • Serve as the technical lead for Halozyme’s Microsoft Fabric-based data lakehouse — workspace topology, OneLake storage design, medallion-layer standards (bronze/silver/gold), Spark and SQL compute configuration, and CI/CD pipeline design
  • Provide technical leadership and mentorship to data and analytics engineers through code reviews, architectural guidance, and hands-on problem solving — personally contributing code and architectural artifacts alongside the team
  • Design and build production-grade data ingestion pipelines across 50+ source systems (ERP, Veeva, LIMS, ELN, CTMS, HRIS, and others) using Data Factory, dataflows, shortcuts, and mirroring patterns
  • Serve as the hands-on technical counterpart to external implementation partners — reviewing deliverables line by line, enforcing engineering standards, and catching quality issues before they reach production
  • Partner directly with business stakeholders across R&D, Commercial, Manufacturing, Business Development, Supply Chain and Corporate functions to understand data requirements, prioritize platform capabilities, and ensure the UDP delivers actionable value aligned to each function’s strategic needs
  • Design and operationalize the platform support model — including intake workflows, tiered support processes, issue triage, and escalation paths — to ensure the data platform remains responsive, reliable, and well-governed as enterprise adoption scales
  • Help define, instrument, and report on key technical KPIs — pipeline reliability, data freshness, query performance, onboarding velocity, and incident resolution time — to drive continuous improvement and demonstrate platform health to leadership
  • Implement and maintain data quality monitoring, alerting frameworks, incident response procedures, and platform SLAs — owning reliability as a commitment, not just a team metric
  • Build and maintain the semantic layer (Power BI datasets, gold-layer views, and self-service analytics capabilities) to ensure business consumers get trusted, performant data products
  • Execute data classification (GxP vs. non-GxP), role-based access control, sensitivity labeling, and audit logging in coordination with IT Security and the Data Governance Lead
  • Technically scope, sequence, and deliver new source system integrations from discovery through production, managing dependencies and communicating timelines with precision
  • Establish and enforce platform engineering standards: naming conventions, branching strategies, testing patterns, documentation requirements, and operational runbooks
  • Other duties as assigned

Benefits

  • Employee Stock Purchase Program
  • 401(k) matching
  • Opportunities to grow in a culture that prioritizes learning, development and progression through in-house programs
  • Tuition reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service