Databricks Data Engineer - India

Cogniify,
Remote

About The Position

We are looking for an experienced Databricks Data Engineer to support, maintain, and enhance existing Databricks-based data applications and pipelines. The role focuses on ensuring reliability, performance, and scalability of production Databricks workloads rather than building net-new platforms from scratch. You will work closely with data, analytics, and engineering teams to keep critical data applications stable, optimized, and aligned with business needs.

Requirements

  • 6–9 years of overall experience in data engineering, with strong hands-on experience in Databricks.
  • Solid proficiency in Apache Spark (PySpark and/or Scala) and SQL.
  • Proven experience supporting and optimizing production Databricks workloads (jobs, notebooks, Delta Lake, workflows).
  • Strong understanding of Delta Lake concepts (ACID transactions, time travel, optimization techniques such as Z-ordering, vacuum, optimize).
  • Experience with Databricks Job clusters, Interactive clusters, and performance tuning (partitioning, caching, shuffle optimization, autoscaling).
  • Familiarity with data modeling, ETL/ELT patterns, and production data pipeline support.
  • Experience working with cloud platforms (preferably Azure, AWS, or GCP) in the context of Databricks.
  • Ability to troubleshoot complex Spark and Databricks issues independently.
  • Strong communication skills and ability to work effectively in a remote, EST-aligned team.

Nice To Haves

  • Experience with Unity Catalog, Databricks SQL, or Lakehouse architecture.
  • Knowledge of CI/CD practices for Databricks (e.g., Databricks Asset Bundles, Git integration, Terraform/ARM templates).
  • Familiarity with orchestration tools (Airflow, Azure Data Factory, or Databricks Workflows).
  • Exposure to data quality frameworks, monitoring tools, or cost optimization initiatives on Databricks.
  • Experience supporting analytics or BI teams consuming Databricks data products.

Responsibilities

  • Support and maintain existing Databricks applications, notebooks, jobs, and Delta Lake pipelines in production.
  • Monitor, troubleshoot, and resolve issues related to job failures, performance degradation, data quality, and cluster utilization.
  • Optimize existing Spark jobs, SQL queries, and Delta tables for cost, performance, and reliability.
  • Manage and improve Databricks workspace configurations, including clusters, job scheduling, access controls, and Unity Catalog (where applicable).
  • Implement and maintain data quality checks, logging, alerting, and basic observability for Databricks workloads.
  • Collaborate with stakeholders to understand requirements for enhancements or bug fixes on existing applications.
  • Perform incremental improvements, refactoring, and technical debt reduction on current Databricks solutions.
  • Ensure adherence to best practices around security, governance, and cost management within the Databricks environment.
  • Document existing pipelines, dependencies, and operational runbooks.
  • Participate in on-call or support rotations as needed to maintain production stability (within EST working hours).
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service