Sr. Databricks Engineer

Concentrix•Usa, TX
•Remote

About The Position

The Senior Data Engineer is responsible for the end-to-end data lifecycle, including the analysis, design, and implementation of data solutions on the Databricks Data Intelligence Platform running on Microsoft Azure, AWS or GCP. The role leverages Databricks, Python, PySpark, and SQL to deliver data ingestion, transformation, modeling, and pipeline orchestration. As the senior engineer, this role leads solution design, sets engineering standards for how pipelines are built, governed, and deployed, and guides other data engineers.

Requirements

  • Experience 10+ years of professional experience in data engineering, data & analytics, or data warehousing.
  • 5+ years of hands-on experience working with Databricks in a data engineering capacity, including production workloads.
  • 5+ years of experience with Python, including manipulating large datasets using PySpark DataFrames, Spark SQL, and/or the pandas API on Spark.
  • 3+ years of data engineering experience on Microsoft Azure (e.g., ADLS Gen2, Azure Data Factory, Key Vault, Microsoft Entra ID).
  • Expert proficiency in Python and SQL for data engineering, data modeling, transformation, and querying.
  • Deep knowledge of Delta Lake and Spark performance tuning, including liquid clustering, OPTIMIZE/VACUUM, predictive optimization, partitioning, data skew, and Photon.
  • Hands-on experience with Unity Catalog for data governance, security, and lineage.
  • Experience implementing streaming data ingestion in Databricks using Structured Streaming, Auto Loader, and Event Hubs or Kafka.
  • Strong understanding of testing and deployment strategies for data pipelines, including unit testing, CI/CD, and productionizing ETL pipelines and Databricks SQL queries.
  • Experience enabling BI consumption through Databricks SQL warehouses and tools such as Power BI.
  • Bachelor’s degree in computer science, Information Systems, Data Engineering, or a closely related field, or equivalent professional experience.

Nice To Haves

  • Active Databricks Certified Data Engineer Professional certification.
  • Active Databricks Certified Data Engineer Associate certification.
  • Experience with healthcare data (e.g., claims, EHR/clinical data, HL7/FHIR) and HIPAA/PHI handling.
  • Infrastructure as code experience with Terraform (Databricks and Azure providers).
  • Experience preparing data for machine learning and generative AI use cases.
  • Familiarity with Delta Sharing and Lakehouse Federation.

Responsibilities

  • Design, develop, and maintain robust batch and streaming data pipelines for extracting, transforming, and loading (ETL/ELT) data from multiple sources into the Databricks lakehouse on Azure.
  • Automate and streamline data ingestion using Auto Loader, Lakeflow Connect, and Azure services such as Azure Data Factory and Event Hubs.
  • Build and manage pipelines with Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables), including data quality expectations and change data capture.
  • Perform complex data manipulation and transformation using PySpark and Spark SQL, ensuring data is prepared for advanced analytics, reporting, and AI/ML use cases.
  • Design and implement orchestration and scheduling of data workflows with Lakeflow Jobs (formerly Databricks Workflows) to ensure reliability and efficiency.
  • Design and implement scalable data models within the lakehouse using medallion architecture (bronze/silver/gold) and dimensional modeling, including SCD Type 1 and Type 2.
  • Implement data governance and security in Unity Catalog, including access controls, lineage, and row filters and column masks for sensitive data.
  • Optimize performance and cost across Delta tables, Spark workloads, serverless and classic compute, and SQL warehouses.
  • Establish CI/CD and testing practices for data pipelines using Git folders, Declarative Automation Bundles (formerly Databricks Asset Bundles), and Azure DevOps or GitHub Actions.
  • Monitor pipeline health and data quality using system tables and data quality monitoring; lead root-cause analysis and resolution of production issues.
  • Lead design and code reviews, define engineering standards, and mentor data engineers on the team.
  • Collaborate with cross-functional teams, including business stakeholders and IT, to ensure data solutions meet business and performance requirements aligned to project goals.

Benefits

  • medical, dental, and vision insurance
  • comprehensive employee assistance program
  • 401(k) retirement plan
  • paid time off and holidays
  • paid learning days
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service