Senior Data Engineer

Marvik•Colorado Springs, CO

About The Position

This is a hands-on building role where you will turn raw, messy fabrication data into clean, well-modeled, AI-ready datasets that our AI/ML and analytics workloads run on. You will be responsible for building and operating scalable ingestion, ELT/ETL, and orchestration pipelines, designing and implementing low-latency, real-time data ingestion flows, and implementing layered architectures using PySpark/SQL. You will also apply data quality and governance techniques, deliver feature-ready datasets, establish testing and monitoring for pipeline observability, and utilize AI-assisted development tools.

Requirements

  • 5+ years of hands-on data engineering experience building and operating production data pipelines at scale.
  • Strong proficiency in Python, SQL, and PySpark / Apache Spark, backed by solid software engineering fundamentals (Git, CI/CD, unit/integration testing).
  • Demonstrated hands-on experience implementing real-time data ingestion and streaming pipelines (not limited to batch processing).
  • Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture).
  • Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred.
  • Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement.
  • Active experience leveraging AI-assisted development tools (Cursor, Copilot, Claude) to accelerate engineering velocity.

Nice To Haves

  • Hands-on experience with Microsoft Fabric (Fabric Lakehouse, Data Factory, Synapse Analytics).
  • Experience extracting data from document stores / NoSQL databases (specifically MongoDB / MongoDB Atlas and Change Streams / CDC).
  • Streaming frameworks experience (Event Hubs, Kafka, Spark Structured Streaming).
  • Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads.
  • AEC / Construction / MEP domain experience.

Responsibilities

  • Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines (batch and real-time streaming) within Microsoft Fabric and cloud lakehouse environments.
  • Design and implement low-latency, real-time data ingestion flows to support live operational analytics and streaming workloads.
  • Implement layered (medallion-style: Bronze/Silver/Gold) architectures using PySpark/SQL with idempotent, backfillable, and incrementally loaded jobs.
  • Apply deduplication, normalization, schema validation, and lineage tracking to ensure downstream data is high-quality, trustworthy, and audit-ready.
  • Deliver feature-ready, curated datasets to support business intelligence, analytics, vector search, and AI/ML agentic workloads.
  • Establish testing, monitoring, and pipeline observability (freshness, volume, schema drift) with clear alerting to resolve failures proactively.
  • Utilize AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for writing pipelines, query tuning, and data transformation scripts.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service