Data Engineer (Azure Databricks + Java + Spark)

American IT SystemsAlpharetta, GA
Onsite

About The Position

We are seeking a skilled Data Engineer with expertise in Azure Databricks, Java, and Spark. The role involves building and optimizing data pipelines, managing data models, and ensuring data governance and security within the Azure cloud environment. This position requires a strong understanding of big data technologies and software engineering best practices.

Requirements

  • 5+ years building batch and/or streaming pipelines with Spark.
  • Deep understanding of Spark fundamentals: lazy evaluation, transformations vs. actions, partitions, shuffles, joins, and caching.
  • Experience with Databricks platform on Azure, including Unity Catalog (catalogues/schemas/tables, permissions, external locations, storage credentials, lineage) and Databricks Jobs/Workflows (task orchestration, cluster policies, retries, parameters).
  • Experience with Git integration (Repos), CI/CD, and environment promotion.
  • Experience with Azure cloud storage and access, including secret management via Databricks secret scopes and Key Vault integration.
  • Experience with Delta Lake features: ACID transactions, time travel, MERGE/UPDATE/DELETE, schema enforcement and evolution.
  • Proficiency in complex SQL with CTEs and window functions.
  • Experience with data modeling and architecture, including Medallion (bronze/silver/gold) patterns for data lakes, partitioning, and file layout.
  • Experience with large scale data processing, including cluster sizing, autoscaling, and cost optimization (spot, pools).
  • Experience with observability tools like Spark UI, metrics, event logs, and Delta logs.
  • Experience with software engineering practices: Python packaging, code quality, testing (pytest, chispa), code reviews.
  • Experience with CI/CD for Databricks (Repos, asset bundles/wheels), artifact versioning.
  • Experience with documentation and stakeholder communication.

Nice To Haves

  • Performance tuning in Spark: broadcast joins, AQE, skew mitigation, shuffle optimization.
  • Structured Streaming and Auto Loader for incremental ingestion.
  • Declarative pipelines with Delta Live Tables (DLT): expectations, CDC, quality rules, event logs.
  • ADLS Gen2: ABFS access, ACLs vs. RBAC, mounts vs. direct access, OAuth/ service principals/managed identity.
  • Delta Lake features: OPTIMIZE, Z ORDER, VACUUM and retention policies, compaction and small files mitigation strategies.
  • Query optimization and execution plan analysis in Databricks SQL.
  • Lakehouse design trade-offs and table constraints.
  • Handling TB scale datasets, throughput vs. latency trade offs.
  • Privilege model in Unity Catalog; table, row, and column level security patterns.
  • PII handling, masking/anonymization, auditability, lineage.
  • Scala, dbt, Kafka/Event Hubs, ADF/Airflow, Terraform, Power BI, data mesh concepts, SCD patterns.

Responsibilities

  • Build and maintain batch and streaming data pipelines using Spark and PySpark.
  • Develop and manage data solutions on the Databricks platform on Azure, including Unity Catalog and Databricks Jobs/Workflows.
  • Implement data processing solutions using Delta Lake, leveraging features like ACID transactions, time travel, and schema evolution.
  • Design and implement data models following Medallion architecture patterns (bronze/silver/gold).
  • Optimize large-scale data processing, including cluster sizing, autoscaling, and cost management.
  • Ensure data governance, security, and compliance, including PII handling and access control.
  • Apply software engineering practices such as Python packaging, code quality, testing, and CI/CD for Databricks.
  • Communicate effectively with stakeholders and document solutions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service