Staff Data Engineer

Harbor Compliance
$172,000 - $215,000

About The Position

Harbor Compliance is building its first end-to-end data platform — unifying fragmented data from HubSpot, financial systems, and HRIS into a single, AI-augmented source of truth for executive decision-making. As Staff Data Engineer, you'll be a foundational technical hire, working closely with the Sr. Director of Data Platform & Analytics to design and build the real-time, AI-ready data infrastructure that powers this platform. This is a hands-on role for someone who wants to build from the ground up, not maintain what already exists.

Requirements

  • 7+ years of hands-on data engineering experience, with meaningful depth in streaming/event-driven systems, not just batch pipelines.
  • Proven experience designing and building near real-time pipelines from scratch (e.g., Kafka, Kinesis, Flink, Debezium/CDC) in a production environment.
  • Hands-on production experience with vector databases and embeddings (e.g., Zilliz, Pinecone, Weaviate, pgvector, Milvus) — ideally having built this infrastructure from the ground up rather than just consuming a managed AI feature.
  • Advanced proficiency in SQL and Python.
  • Working knowledge of a cloud warehouse/lakehouse platform (Snowflake, BigQuery, or Databricks) and dbt — you'll use these, but they're the storage/transform layer supporting the streaming and AI work, not the main focus.
  • Proven experience building or materially contributing to an end-to-end production data environment, ideally as an early or founding data hire.
  • Familiarity with B2B SaaS and recurring revenue data models (customer lifecycle, pipeline/conversion data).
  • Working knowledge of BI/reporting tools (e.g., Looker, Tableau, Power BI) as a downstream consumer of your data models.
  • Ability to work independently and drive multi-stakeholder projects in a lean, scrappy, fast-moving environment — comfortable with ambiguity and building without a lot of existing infrastructure or process.

Nice To Haves

  • Strong command of event streaming and CDC tooling.
  • Hands-on experience with vector databases (Pinecone, Weaviate, pgvector, Milvus, or similar), including embedding strategies and chunking approaches for retrieval use cases.
  • Working knowledge of ELT/ETL tools (e.g., Fivetran, Airbyte) for batch use cases.
  • Working knowledge of data observability/reliability practices — automated alerting, lineage tracking.
  • Fluency with AI-augmented development tools (e.g., Claude, Copilot) to accelerate engineering and documentation.

Responsibilities

  • Design, build, and own near real-time data pipelines (CDC, streaming ingestion, event-driven architectures) as the backbone of the platform's data flow
  • Evaluate, implement, and maintain vector database infrastructure and embedding pipelines to support AI-augmented use cases (semantic search, retrieval-augmented generation, AI agents acting on company data).
  • Design and build ELT/ETL pipelines ingesting data from our platform, financial platforms, CRM (HubSpot) and HRIS, feeding both real-time and batch use cases.
  • Partner with the Sr. Director to architect the underlying warehouse/lakehouse as a supporting system of record — the storage layer beneath the streaming and AI infrastructure.
  • Build lightweight transformation layers (e.g., dbt) as needed to enable our Analytics Engineering team translate raw data into business-ready datasets aligned to core metrics like ARR, CAC, and churn.
  • Own pipeline reliability and observability — monitoring, automated failure alerting, and lineage tracking across both streaming and batch pipelines.
  • Build the technical foundation for self-service and AI-powered reporting, partnering with BI, Product & Engineering stakeholders on recurring executive and departmental reports.
  • Implement data governance practices, including documentation standards and access controls.
  • Partner cross-functionally with Finance, Marketing, Customer Success, and Operations to translate data needs into reliable, low-latency data products.
  • Leverage AI-augmented development workflows (e.g., Claude Code) to accelerate pipeline development and documentation.

Benefits

  • health benefits
  • flexible paid time off
  • parental leave
  • fertility and adoption assistance
  • 401(k)
  • educational reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service