Data Engineer

Kumaran SystemsToronto, ON

About The Position

We are seeking a skilled Data Engineer to join our team. The ideal candidate will have strong hands-on experience with DataBricks and Apache Spark, along with a solid understanding of SQL and data transformation techniques. You will be responsible for developing and maintaining ETL tools and data pipelines, working with cloud platforms like Azure, AWS, or GCP, with a strong preference for Azure. A good grasp of data warehousing concepts, data modeling, and performance tuning in Spark is essential. Experience with CI/CD pipelines, DevOps practices, data governance, and security is also required. This role involves extensive work with Azure DataBricks, Delta Lake, Unity Catalog, and building various types of data pipelines including batch (autoloader) and Spark structured streaming. You will also be involved in creating end-to-end environments, managing data models like SCDs and CDCs, and utilizing Lakehouse federation.

Requirements

  • Strong hands-on experience with DataBricks and Apache Spark (PySpark/Scala).
  • Experience in SQL and data transformation techniques.
  • Knowledge of ETL tools and data pipeline development.
  • Experience working with cloud platforms (Azure/AWS/GCP).
  • Strong Azure cloud background.
  • Understanding of data warehousing concepts.
  • Strong problem-solving and analytical skills.
  • Hands-on experience with Azure Data bricks or Delta Lake, in building ETL pipelines: batch (autoloader) and Spark structured streaming.
  • Knowledge of data modelling and performance tuning in Spark.
  • Familiarity with data governance and security practices.
  • Strong hands-on working experience of Unity catalog.
  • Hands-on exposure to creating end-to-end environments: creating catalogs, schemas, tables, materialized views, functions, volumes.
  • Experience in building SCD 1 and SCD 2 (slowly changing dimensions) on dimension tables.
  • Experience in building CDC (change data capture pipelines).
  • Strong hands-on experience with Lakehouse federation, creating foreign catalogs to get data from external sources.
  • Strong understanding of Databricks partitioning and Liquid clustering.

Nice To Haves

  • Exposure to CI/CD pipelines and DevOps practices.

Responsibilities

  • Build and maintain ETL tools and data pipelines.
  • Develop batch (autoloader) and Spark structured streaming pipelines using Azure DataBricks or Delta Lake.
  • Implement SCD 1 and SCD 2 on dimension tables.
  • Build CDC (change data capture) pipelines.
  • Utilize Unity Catalog for creating catalogs, schemas, tables, materialized views, functions, and volumes.
  • Implement Lakehouse federation by creating foreign catalogs to access external data sources.
  • Optimize Spark performance through data modeling and tuning.
  • Ensure data governance and security practices are followed.
  • Integrate CI/CD pipelines and DevOps practices into the data engineering workflow.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service