AI Data Engineer, Data Platform

Collective•San Francisco, CA
•$180,000 - $230,000•Hybrid

About The Position

Collective is seeking a Data Engineer to own and scale the data platform that powers analytics, reporting, and AI across the company. This role involves designing, building, and maintaining data pipelines, modeling data into reliable tables, and establishing engineering standards for the platform. The Data Engineer will work within the Data Engineering team, collaborating with product engineers, analysts, and business stakeholders. This is a hands-on position for an individual who prioritizes data quality, takes end-to-end ownership of production systems, and aims to build the foundational data infrastructure for the company.

Requirements

  • 5+ years of professional experience in data engineering, analytics engineering, or a closely related role, ideally at a B2B SaaS or fintech company.
  • Expert-level SQL and strong Python skills for building pipelines, transformations, and tooling; comfortable writing tested, production-grade code.
  • Hands-on production experience with a cloud data warehouse (BigQuery strongly preferred), dbt or equivalent transformation framework, managed ingestion tools (Fivetran or similar), and an orchestrator (Airflow, Dagster, Cloud Composer, or similar).
  • Deep understanding of dimensional modeling, layered warehouse architecture, and schema design, with strong opinions on grain, naming, and consistency.
  • Experience implementing testing frameworks, lineage, monitoring, and alerting for data pipelines, and operating them in production including on-call.
  • Fluency with git-based workflows, code review, CI/CD, and infrastructure-as-code; treating data infrastructure as software.
  • A track record of taking ambiguous, high-impact problems and delivering reliable systems end-to-end, with a focus on outcomes rather than just implementation.
  • Ability to explain technical trade-offs to non-technical stakeholders and drive alignment on data definitions across teams.

Nice To Haves

  • Experience with streaming or event data (Pub/Sub, Kafka, or similar) and product analytics tooling (Amplitude or similar).
  • Experience with Terraform and Google Cloud Platform infrastructure.
  • Experience with observability platforms such as Datadog.
  • Exposure to financial, accounting, tax, or payroll data and the correctness requirements that come with it.
  • Experience building semantic layers or metric stores consumed by LLM-based tools, or supporting LLM evaluation programs.
  • AI-assisted development experience (Claude Code or similar).

Responsibilities

  • Design and build data pipelines, including developing, deploying, and maintaining scalable batch and event-driven pipelines that ingest data from application databases, SaaS tools, and external APIs into BigQuery using managed connectors (Fivetran), custom Python loaders, and orchestration tooling.
  • Model the data by designing and implementing dimensional and analytical data models in dbt, following a layered architecture (raw, staging, marts) with clear grain, naming conventions, and documentation.
  • Own data quality and reliability by implementing testing, monitoring, alerting, and data contracts across the pipeline; defining and meeting freshness and accuracy SLAs; and triaging and resolving pipeline failures and data incidents.
  • Optimize performance and cost by tuning warehouse queries, partitioning, and clustering; managing BigQuery spend; and keeping pipelines efficient as data volume grows.
  • Establish engineering standards by driving best practices for version control, code review, CI/CD, and infrastructure-as-code across the data stack; documenting systems and runbooks.
  • Govern and secure data by implementing access controls, PII handling, and data retention practices appropriate for a financial services company; partnering with Security and Legal on compliance requirements.
  • Enable the business by partnering with product engineers on source schema design and change management, and with analysts and stakeholders to translate business questions into reliable datasets, metric definitions, and self-serve reporting in Metabase.
  • Support AI and analytics use cases by maintaining the semantic layer, metric definitions, and documentation that allow LLM-based tools and internal agents to query the warehouse accurately and consistently.

Benefits

  • Hybrid Work Model
  • Fresh Lunch provided on in-office days
  • $150 monthly reimbursement for transit expenses
  • $200 quarterly reimbursement for well-being
  • Flexible PTO plus 14 company holidays
  • 100% medical, dental, and vision for employees; 75% coverage for dependents
  • 16 weeks fully paid parental leave
  • 401k plan
  • Equity package
  • Quarterly virtual events and an annual in-person summit
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service