Data Engineer

Spectrum Health CareToronto, ON
CA$95,000 - CA$105,000

About The Position

Spectrum Health Care (Spectrum) is building a centralized Cloud Data Lakehouse to support clinical analytics, AI initiatives, and business intelligence across the organization. The Data Engineer will play a key role in Spectrum’s digital transformation and innovation, supporting high-impact initiatives to migrate data from legacy EMR systems, cutting-edge EHR platforms, and critical business applications into our cloud stack ecosystem and Salesforce Health Cloud. In this role, you will design, build, and optimize batch and real-time data pipelines, develop OLTP and OLAP data models, and leverage enterprise-grade ETL tools such as Informatica to deliver secure, compliant, and high-quality datasets. Your work will provide the trusted data foundation to support clinical decision-making, operational excellence, analytics, and AI-drive initiatives across our organization. This role is for individuals passionate about data engineering, cloud technologies, and delivering high-quality data that drives meaningful outcomes for patients, families and the communities we proudly serve.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, Data Analytics, or a related technical field
  • 3-7 years of hands-on data engineering experience, with a proven track record of executing complex data migration projects and building cloud data lake architectures.
  • Expertise in ER (Entity-Relationship) diagramming, relational database modeling (3NF for OLTP), and dimensional modeling (Star Schema, Snowflake, Medallion) for OLAP systems
  • Experience with enterprise ETL platforms such as Informatica (Informatica PowerCenter / IDMC), Azure Data Factory, SSIS, or Talend for legacy data extraction and transformation
  • Experience with cloud data architectures (Azure data services: Azure Data Factory, ADLS Gen2, Azure Synapse, Databricks) and familiarity with streaming/real-time pipelines (Event Hubs, Kafka, Spark Streaming)
  • Experience working with Salesforce data structures, SOQL, Salesforce Bulk API, and data loading tools (e.g., Informatica Salesforce Data Loader, MuleSoft, Fivetran, dbt, or custom Python scripts)
  • Expert-level SQL skills (complex joins, CTEs, window functions, schema design, index tuning, and database optimization)
  • Strong Python skills for data manipulation, scripting, and pipeline execution using libraries like Pandas, PySpark, and SQLAlchemy

Responsibilities

  • Design, deploy and maintain a centralized Azure ADLS Gen2/Databricks/Synapse (or Snowflake) data lake that serves as the single source of truth
  • Create batch & real‑time ETL/ELT flows (Azure Event Hubs, Kafka, Spark Streaming) to ingest EHR data and operational telemetry
  • Write and run complex migration scripts using Informatica PowerCenter/IDMC or Azure Data Factory to move data from Epic, other EHRs and legacy OLTP stores into the lakehouse and Salesforce Health Cloud
  • Build automated validation, deduplication, masking, encryption and column‑level security to guarantee 100% integrity and PHI compliance
  • Build high‑throughput pipelines (SOQL, Bulk API 2.0, Informatica, MuleSoft, dbt, Python) that sync Health Cloud data with the central repository
  • Transform HL7, FHIR and EDI messages into clean canonical schemas for storage and analytics
  • Design Robust Data Models – Perform ER/3NF modeling for OLTP and dimensional (star/snowflake/medallion) modeling for OLAP analytics
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service