Data Engineer

Cynet SystemsAlpharetta, GA

About The Position

We are seeking a skilled Data Engineer with extensive experience in building and maintaining robust data pipelines. The ideal candidate will have a deep understanding of Spark fundamentals, experience with the Databricks platform on Azure, and strong SQL and data modeling skills. This role involves designing, implementing, and optimizing data solutions, ensuring data governance, security, and compliance, and collaborating with stakeholders to deliver technical documentation and support data initiatives.

Requirements

  • 5+ years building batch and/or streaming pipelines with Spark.
  • Deep understanding of Spark fundamentals including lazy evaluation, transformations vs. actions, partitions, shuffles, joins, and caching.
  • Experience with Databricks platform on Azure including Unity Catalog (catalogs/schemas/tables, permissions, external locations, storage credentials, lineage).
  • Proficiency with Databricks Jobs/Workflows for task orchestration, cluster policies, retries, and parameters.
  • Experience with Git integration (Repos), CI/CD, and environment promotion.
  • Knowledge of cloud storage and access on Azure, specifically secret management via Databricks secret scopes and Key Vault integration.
  • Experience with Delta Lake including ACID transactions, time travel, MERGE/UPDATE/DELETE, and schema enforcement/evolution.
  • Strong SQL skills including complex SQL with CTEs and window functions.
  • Experience with data modeling and architecture using Medallion (bronze/silver/gold) patterns for data lakes.
  • Ability to design tables for CDC and incremental processing.
  • Experience with large scale data processing including cluster sizing, autoscaling, and cost optimization.
  • Familiarity with observability tools such as Spark UI, metrics, event logs, and Delta logs.
  • Knowledge of software engineering practices including Python packaging, code quality, and testing (pytest, chispa).
  • Strong verbal and written communication skills.
  • 10-12 years of professional experience in data engineering or related fields.
  • Extensive experience with the Azure cloud ecosystem.

Nice To Haves

  • Performance tuning: broadcast joins, AQE, skew mitigation, and shuffle optimization.
  • Structured Streaming and Auto Loader for incremental ingestion.
  • Declarative pipelines with Delta Live Tables (DLT).
  • ADLS Gen2: ABFS access, ACLs vs. RBAC, and managed identity.
  • OPTIMIZE, Z ORDER, VACUUM and retention policies.
  • Compaction and small files mitigation strategies.
  • Query optimization and execution plan analysis in Databricks SQL.
  • Lakehouse design trade-offs and table constraints.
  • Handling TB scale datasets and throughput vs. latency trade-offs.
  • Privilege model in Unity Catalog; table, row, and column level security patterns.
  • PII handling, masking/anonymization, and auditability.
  • Scala, dbt, Kafka/Event Hubs, ADF/Airflow, Terraform, Power BI, data mesh concepts, and SCD patterns.

Responsibilities

  • Build and maintain batch and streaming data pipelines using Spark and PySpark.
  • Manage Databricks platform components including Unity Catalog and Workflows.
  • Implement data modeling strategies such as partitioning and file layout to match read patterns.
  • Ensure data governance, security, and compliance within the Databricks environment.
  • Execute CI/CD processes for Databricks using Repos and asset bundles.
  • Collaborate with stakeholders and provide technical documentation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service