Founding Data Engineer

HollyNew York, NY
$170,000 - $216,000Onsite

About The Position

We're hiring our Founding Data Engineer to build the data backbone of Holly. Local governments publish enormous amounts of public information — salary schedules, job classifications, MOUs, budgets — but it's scattered across thousands of websites and buried in messy formats: scanned PDFs, inconsistent HTML, spreadsheets, and everything in between. Turning that chaos into clean, trustworthy, structured data is our single biggest data challenge and one of our deepest moats. You'll design and build the data platform: the pipelines and systems that ingest public government data at scale, normalize and validate it, and serve it as clean, canonical datasets the rest of Holly's product builds on as its source of truth. Over time, you'll make this platform increasingly automated and intelligent — less manual wrangling, more self-healing, monitored, high-quality pipelines. This is a hands-on, high-ownership building role. As the first data hire on a small engineering team, you'll set the direction, make the architecture calls, and establish the standards for how Holly collects, models, and trusts its data for years to come. You'll work directly with the founders and partner closely with our product engineers - your job is to make sure they always have the clean, reliable data they need to build features. If you've worked with large, high-volume data and love the challenge of taming messy real-world inputs into something people can rely on, we'd love to talk.

Requirements

  • Are a senior data engineer with a strong track record building and operating production data systems (several years of relevant experience or equivalent)
  • Have worked with large-scale, high-volume data — ideally where lots of sources, users, or records make volume and reliability matter
  • Are strong at data modeling and SQL, with experience designing schemas that others build on (Postgres a plus)
  • Have built and owned ETL/ELT pipelines that handle messy, heterogeneous, real-world inputs (scraped data, PDFs, HTML, spreadsheets)
  • Bring a strong data-quality mindset — validation, testing, monitoring, lineage, and reliability are core to how you work
  • Take ownership and move fast — you work independently, ship often, and thrive in early-stage ambiguity
  • Have a growth mindset — you learn quickly, and raise the bar through collaboration and clear standards
  • Are pragmatic about tooling and comfortable working in (or ramping quickly into) a modern TypeScript/Postgres codebase

Nice To Haves

  • Experience with large-scale web scraping / crawling, document extraction (OCR), or LLM-assisted parsing
  • Experience with embeddings / vector search or supporting ML/AI data workflows
  • Experience with analytical/columnar or warehouse stacks (ClickHouse, BigQuery, Snowflake) and/or streaming pipelines
  • Comfort in TypeScript/Node (our stack)
  • Experience in government, public sector, or civic tech
  • Prior early-stage startup experience

Responsibilities

  • Own and Build the Data Platform: Own the data platform end-to-end — from raw public sources to clean, canonical datasets the product consumes
  • Design the architecture, schemas, and standards for how Holly ingests, models, and trusts its data
  • Partner with the founders to scope work, make tradeoffs, and drive delivery on our highest-leverage data initiatives
  • Set the long-term direction for our data foundation as the first data hire
  • Build Ingestion & Normalization Pipelines: Build systems that collect large volumes of public government data from thousands of local-government sources across the web
  • Turn messy, heterogeneous inputs — scanned PDFs, inconsistent HTML, spreadsheets — into structured, normalized data (parsing, extraction, OCR, dedupe, entity resolution, schema mapping)
  • Where it adds leverage, incorporate LLM-assisted extraction and embeddings into the pipeline
  • Build for freshness, reliability, and scale so data stays current and trustworthy
  • Model & Serve Data for the Product: Design canonical data models and domain schemas that product engineers build on
  • Expose clean, versioned, well-documented datasets the main app can reliably consume
  • Own data quality, validation, lineage, and observability so downstream teams can trust what they're building on
  • Make It Automated & Intelligent: Evolve pipelines from manual/one-off toward automated, self-healing, monitored systems
  • Establish data-quality checks, alerting, and standards that keep the platform reliable as it grows
  • Raise the bar on how we collect, validate, and serve data across the company

Benefits

  • Foundational ownership — Architect and build the data platform that will define our product and data model for years.
  • Technical influence — Make the high-leverage calls on data architecture, standards, and how we scale.
  • High-impact scope — Build the data foundation the entire product depends on, and see it power features customers rely on quickly.
  • Public-service impact — Your work improves how local governments operate, helping millions of Americans access public-service careers.
  • Competitive package — $170K - $216K base, 0.15–0.4% equity (L3), comprehensive health benefits (platinum plan, vision, dental), 401(k), paid parental leave, and a professional development stipend.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service