Data Engineer

SMBCNew York City, NY
Hybrid

About The Position

We are seeking a Databricks Engineer with AWS expertise to build, optimize, and maintain our enterprise data Lakehouse infrastructure. In this role, you will be responsible for designing high-performance data pipelines, implementing regulatory and real time reporting along with advanced analytics environments, and ensuring seamless integration between Databricks and core AWS services to support real-time financial trading data consumption.

Requirements

  • Professional experience in architectural design and development within the Databricks platform, working in an AWS cloud environment.
  • Proven proficiency in automated deployments using Databricks Asset Bundles (DABs), Terraform (specifically the Databricks and AWS providers), and standard Git pipelines (e.g., GitHub Actions, GitLab CI/CD, or AWS CodePipeline).
  • Programming skills in Python (PySpark) and SQL for complex data manipulation and transformation.
  • Strong understanding of Apache Spark internals, Delta Lake mechanics, and streaming data concepts (e.g., interacting with Amazon MSK or Kafka data streams).
  • Proven experience building production-grade ETL/ELT pipelines, handling data schema validation, and cleansing raw capture feeds.

Nice To Haves

  • Databricks Certified Data Engineer Professional or AWS Certified Data Engineer – Professional is highly advantageous.

Responsibilities

  • Design and implement robust data pipelines using the Databricks Medallion Architecture (Bronze, Silver, Gold layers) to process structured and unstructured data.
  • Develop, scale, and orchestrate complex data workflows utilizing Databricks Jobs and Delta Live Tables (DLT).
  • Standardize and automate the deployment of Databricks assets, workspace configurations, and code pipelines across Dev, QA, and Production environments.
  • Ensure seamless data cataloging, storage, and movement across the AWS ecosystem, specifically integrating Databricks with Amazon S3, AWS Glue, and AWS IAM for secure access control.
  • Tune Spark clusters, optimize Delta Lake storage (e.g., Z-Ordering, partitioning), and manage compute costs within the AWS environment.
  • Implement fine-grained data access controls, data lineage, and auditing using Unity Catalog or native cloud security controls.
  • Partner with Data Scientists, Risk Managers, and downstream analytics teams to deliver clean, business-ready data views for reporting and AI modeling.

Benefits

  • Competitive portfolio of benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service