Senior Data Engineer

Jellyfish
$165,000 - $235,000

About The Position

Jellyfish processes a huge amount of engineering data, and we are investing heavily in the foundations that make that data reliable, governable, and easy to use. We are looking for a Data Engineer to help mature our Databricks-based data platform, establish strong data modeling patterns, and build the systems that move data from raw ingestion to trusted production datasets. You’ll work across ingestion, transformation, storage, governance, and serving. If you enjoy turning messy data pipelines into durable platform architecture and want to help define how a modern lakehouse should actually operate, you’re the perfect fit.

Requirements

  • Extensive experience with Databricks, Spark, Delta Lake, or a comparable lakehouse platform and understand how to operate it beyond simply writing notebooks.
  • Understanding of partitioning, incremental processing, schema evolution, distributed execution, file formats, and the performance characteristics of large analytical datasets.
  • Ability to reason about canonical entities, relationships, grain, dimensional modeling, and the boundary between platform models and consumer-specific models.
  • Design pipelines to be observable, retryable, idempotent, and understandable when they fail.
  • Understanding of how object storage, compute, networking, IAM, and managed data services fit together in a modern cloud data architecture.
  • Pragmatic platform builder who cares about standards and architecture, but also knows when to ship a practical solution and iterate.
  • Humility, a performance-driven attitude, and a team-player approach.
  • Passion for building great companies in an environment where a sense of humor is a must.
  • Must be authorized to work for any employer in the US. Unable to sponsor or take over sponsorship of an employment visa at this time.

Nice To Haves

  • Helped build or migrate to a medallion-style lakehouse architecture.
  • Worked with Databricks Unity Catalog, OpenMetadata, or another governance and lineage platform.
  • Implemented CDC pipelines from PostgreSQL, RDS, or Aurora.
  • Worked with Airflow or another production workflow orchestration platform.
  • Moved analytical data into serving systems like ClickHouse, Snowflake, BigQuery, or similar platforms.
  • Helped introduce data contracts, canonical schemas, or platform-wide data quality standards.

Responsibilities

  • Build and maintain data pipelines and datasets in Databricks and Delta Lake, improving reliability, performance, and operational visibility across the platform.
  • Help establish clear Bronze, Silver, and Gold layer responsibilities, including standards for schema evolution, transformation ownership, data retention, and promotion between layers.
  • Design durable canonical models for core Jellyfish entities and relationships. Work with application and analytics teams to ensure downstream datasets are structured around consistent definitions rather than one-off transformations.
  • Build and improve batch and incremental pipelines using technologies like Databricks, Airflow, Spark, and cloud object storage. Focus on idempotency, scalability, observability, and recoverability.
  • Work with catalog and governance tooling to establish lineage, ownership, schema standards, quality checks, and discoverability across the platform.
  • Help create reliable patterns for moving curated data from Databricks into systems like ClickHouse and other future serving destinations without tightly coupling the platform to any single database.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service