Data Engineer Consultant

Indicium AINew York, NY
$120,000 - $175,000

About The Position

As a Data Engineer Consultant on our US team, you will build and deploy production-grade data pipelines for major enterprise clients. You will spend your day-to-day writing code, modernizing brittle legacy setups into cloud-native Lakehouses, and making client data clean, fast, and ready for AI applications. In this role, you will work directly with modern technologies like Databricks, dbt, Python, and Terraform, collaborating with technical clients and our global engineering squads to ship solutions in weeks, not months.

Requirements

  • Strong background writing production code in Python and advanced SQL.
  • Hands-on experience building on the Databricks Lakehouse platform (Delta Lake, PySpark, Unity Catalog, Workflows).
  • Solid experience using dbt for data transformation, testing, and documentation.
  • Practical experience writing Terraform scripts to deploy AWS or GCP cloud data infrastructure.
  • Proficient with Git (GitHub/GitLab) for CI/CD workflows and experience with schedulers like Airflow.
  • Ability to communicate technical decisions clearly to client teams and collaborate in a fast-paced consulting environment.
  • Ability to jump into unfamiliar legacy codebases, figure out how the data flows, and rebuild it cleanly without needing a rigid playbook.
  • Passion for writing testable, documented, and reusable code.

Nice To Haves

  • Databricks, AWS, or dbt certifications.
  • Experience working on migration projects from legacy warehouses (Teradata, Netezza, Snowflake, Redshift) to Databricks.
  • Familiarity with streaming tools like Apache Kafka or CDC (Change Data Capture) pipelines.

Responsibilities

  • Build Lakehouse Modernization Pipelines: Convert legacy data setups (Informatica, PL/SQL, legacy data warehouses) into high-performance PySpark, Databricks SQL, and Delta Live Tables (DLT) using Medallion Architecture standards (Bronze, Silver, Gold).
  • Data Modeling & Transformation: Write clean, modular, and tested dbt models and advanced SQL to transform raw client data into trusted, business-ready datasets.
  • Automate Infrastructure (DataOps): Use Terraform to provision cloud resources (AWS/GCP) programmatically and enforce data governance using Unity Catalog.
  • Integrate & Orchestrate: Build and monitor automated ELT/ETL workflows using tools like Apache Airflow and Databricks Workflows, ensuring data quality, lineage, and strict SLA compliance.
  • Client Engagement & Collaboration: Embedded directly within client projects, participating in daily standups, code reviews, and working closely with our nearshore delivery squads to deliver on time.
  • Troubleshoot & Optimize: Debug failing pipelines, tune complex SQL/Spark queries for performance, and reduce cloud computing costs for our clients.

Benefits

  • Annual discretionary bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service