Senior Software Engineer, Data & AI

HollyNew York, NY
$170,000 - $216,000Onsite

About The Position

We're hiring a Senior Software Engineer, Data & AI to build Holly's Data Platform. Local governments publish enormous amounts of public information salary schedules, job classifications, MOUs (Labor Union Agreements), budgets but it's scattered across thousands of websites and buried in messy formats: scanned PDFs, inconsistent HTML, spreadsheets, and everything in between. Turning that chaos into clean, searchable, trustworthy data is one of our biggest product challenges and deepest moats. You'll design and build the Data Platform: ingestion and processing pipelines, canonical datasets, semantic search and retrieval, and inference APIs that product engineers can build on. AI is part of the infrastructure, not a layer added later. You'll use models where they improve extraction, classification, normalization, matching, and search, then evaluate and operate those systems in production. This is a hands-on, high-ownership engineering role. You'll own major data systems, make architecture calls, and raise the standards for how Holly collects, models, and trusts its data. You'll work directly with the founders and partner closely with product engineers to make sure they have clean, reliable data to build on. If you've worked with large, high-volume data and love the challenge of taming messy real-world inputs into something people can rely on, we'd love to talk.

Requirements

  • 5+ years building and shipping production software (or equivalent experience)
  • A strong track record building and operating production data systems
  • Worked with large-scale, high-volume data ideally where lots of sources, users, or records make volume and reliability matter
  • Strong at data modeling and SQL, with experience designing schemas that others build on (Postgres a plus)
  • Built and owned ETL/ELT pipelines that handle messy, heterogeneous, real-world inputs (scraped data, PDFs, HTML, spreadsheets)
  • Built production search, retrieval, inference, or AI-assisted data-processing systems
  • Know how to evaluate model quality and operate model-backed systems when outputs are probabilistic
  • Bring a strong data-quality mindset validation, testing, monitoring, lineage, and reliability are core to how you work
  • Take ownership and move fast you work independently, ship often, and thrive in early-stage ambiguity
  • Have a growth mindset you learn quickly, and raise the bar through collaboration and clear standards
  • Pragmatic about tooling and comfortable working in (or ramping quickly into) a modern TypeScript/Postgres codebase

Nice To Haves

  • Open source contributions to or maintainer of a widely used tool.
  • Experience with large-scale web scraping / crawling, document extraction (OCR), or LLM-assisted parsing
  • Experience with embeddings / vector search or supporting ML/AI data workflows
  • Experience with analytical/columnar or warehouse stacks (ClickHouse, BigQuery, Snowflake) and/or streaming pipelines
  • Experience in government, public sector, or civic tech
  • Prior early-stage startup experience

Responsibilities

  • Own and Build the Data Platform end-to-end from raw public sources to clean, canonical datasets the product consumes
  • Design the architecture, schemas, and standards for how Holly ingests, models, and trusts its data
  • Partner with the founders to scope work, make tradeoffs, and drive delivery on our highest-leverage data initiatives
  • Help set the long-term direction for our data foundation
  • Build systems that collect large volumes of public government data from thousands of local-government sources across the web
  • Turn messy, heterogeneous inputs scanned PDFs, inconsistent HTML, spreadsheets into structured, normalized data (parsing, extraction, OCR, dedupe, entity resolution, schema mapping)
  • Where it adds leverage, incorporate LLM-assisted extraction and embeddings into the pipeline
  • Build for freshness, reliability, and scale so data stays current and trustworthy
  • Design canonical data models and domain schemas that product engineers build on
  • Expose clean, versioned, well-documented datasets the main app can reliably consume
  • Own data quality, validation, lineage, and observability so downstream teams can trust what they're building on
  • Build semantic search and retrieval over large, changing government datasets
  • Design inference APIs for extraction, classification, matching, and other AI-powered data processing
  • Build evals, tracing, and fallbacks so model behavior is measurable and dependable
  • Give product engineers clear interfaces for using data and AI capabilities
  • Evolve pipelines from manual/one-off toward automated, self-healing, monitored systems
  • Establish data-quality checks, alerting, and standards that keep the platform reliable as it grows
  • Raise the bar on how we collect, validate, and serve data across the company

Benefits

  • End-to-end ownership. Architect and build data systems that define how the product works.
  • Technical influence. Make the high-leverage calls on data architecture, standards, and how we scale.
  • Commitment to Open Source. We are big believers in supporting open source, and provide a monthly day of Open Source where you can work on your favorite tool. In addition to internal hackathons and other projects
  • Direct access. Work directly with the founders, with autonomy to drive major initiatives end-to-end.
  • High-impact scope. Build the data foundation the entire product depends on, and see its power features customers rely on quickly.
  • Public-service impact. Your work improves how local governments operate, helping millions of Americans access public-service careers.
  • Competitive package. $170k-$216k base, 0.15-0.40% equity (L3), comprehensive health benefits (platinum plan with vision and dental), 401(k), paid parental leave, and a professional development stipend.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service