Data Engineer

Ness Digital Engineering

About The Position

This role focuses on building and managing data pipelines within a Databricks environment, adhering to medallion architecture principles and Unity Catalog governance. The Data Engineer will be responsible for ingesting data into the bronze layer, transforming it into the silver layer, and creating gold marts for analytical purposes. Key tasks include implementing attribution logic, ensuring data quality, managing data volume, and working within existing CI/CD and development workflows.

Requirements

  • Advanced Databricks engineering: Delta Lake, medallion architecture, Databricks Workflows, Auto Loader and incremental ingestion patterns.
  • Unity Catalog to a governance standard — catalogs, schemas, permissions, lineage — not merely as a place tables happen to live.
  • Strong Python and PySpark, and strong SQL. Notebook-based development.
  • Ingestion from REST APIs including pagination, throttling, incremental watermarks and credential handling, plus cloud object storage across AWS, Azure and GCP.
  • Performance and cost optimization of Spark workloads: partitioning, clustering, file sizing and cluster configuration.
  • CI/CD for Databricks — asset bundles or equivalent — and Git-based development workflow.
  • Able to work to an existing catalog structure and coding standard rather than introducing a parallel approach.

Responsibilities

  • Build ingestion into the bronze layer for assigned sources: gateway and observability logs, productivity tool admin APIs, AI-enabled SaaS usage, hyperscaler billing exports and reference data. Land raw and untransformed, on a scheduled refresh, replayable if the downstream design changes.
  • Work to the shared bronze landing contract so each tool is ingested once and serves both this program and the parallel productivity initiative, rather than being integrated twice.
  • Build the silver layer: typed, deduplicated and conformed to the canonical dimensions, refreshed independently of any downstream publication schedule.
  • Build gold marts carrying attribution method, attribution level, cost basis and provisional status alongside cost and usage.
  • Implement the attribution and allocation logic designed by the analysts, including precedence resolution and ratio-based splitting of shared endpoint cost.
  • Work within Unity Catalog governance — shared bronze and silver, separate gold marts with a recorded owner per dataset — including permissions, lineage and cataloging.
  • Implement data quality rules and monitoring: completeness, freshness and tag-coverage checks with alerting, so pipeline problems surface before they reach a divisional invoice.
  • Manage the volume impact of enabling caller-identity data in the cost and usage report, which multiplies row counts by the number of calling identities per model.
  • Work to the per-source cadence — daily where controls and anomaly detection depend on it, monthly where they do not — within the team's existing CI/CD and promotion practices.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service