Senior Software Engineer, Data Engineering

NationGraphToronto, ON
CA$150,000 - CA$200,000Onsite

About The Position

NationGraph is making public sector data accessible and actionable for businesses selling to cities, counties, state agencies, schools, and special districts. NationGraph's data and intelligence engine provides buying signals derived from millions of public sector sources. Founded in 2024, NationGraph is dedicated to making uncommon knowledge common, because public data should actually be public. Learn more at nationgraph.com You'll Join A Team That: Has successfully built, scaled, and sold companies in the past. Built software infrastructure processing billions of dollars in transactions. Is backed by world-class VCs and operating partners who've invested in, and built, iconic companies. About The Role: Own the systems behind NationGraph's unique differentiated asset: the data acquisition platform that collects public information/records at scale, (purchase orders, contracts, various procurement documents) from 100k+ government entities across the country. No one else has solved this government data problem. You'll be one of the people who helps change that. This is real data engineering, not dashboard plumbing. You'll build high-volume interaction and ingestion pipelines with retries, dead-letter handling, checkpointing, and resumability; turn messy government documents (PDFs, spreadsheets, scans, portals) into clean structured data; and keep it flowing into the product reliably. Build with AI at the core: LLM-driven document extraction, classification, and form processing at scale are already in production here, but there's a wide-open opportunity to make them more modular, rigorous, cheaper, and faster. Also design some human-in-the-loop systems. Some pipelines pair automation with a dedicated operations team; you'll contribute to the internal platform they work in every day, complying with access controls, audit logging, and workflow tooling meeting SOC 2 requirements. Own problems end to end in a small, senior team. You'll take a goal like "automate records requests through web portals" from research to design doc to production, and you'll see your work show up directly in customer accounts and revenue. Every decision you make materially changes the trajectory of the company.

Requirements

  • 5+ years building production backend systems
  • Strong TypeScript/Node - the platform is a TypeScript monorepo (API + React frontend). You write well-structured, typed services and know your way around async execution, streams, and queues
  • Pipeline engineering depth - idempotent processing, retries with backoff, dead-letter handling and recovery, checkpoint/resume for long-running jobs, backpressure, graceful degradation when downstream services fail
  • Solid PostgreSQL - schema design, safe migrations on live tables, indexing strategies, query optimization, and comfort operating a database that other teams and the product depend on
  • Document processing at scale - extracting reliable structured data from large volumes of inconsistent files (PDFs, spreadsheets, scans), with deduplication and edge-case handling
  • LLM integration - calling model APIs at scale for extraction and classification, managing rate limits and cost across providers, structuring prompts and outputs so they're testable
  • Third-party API fluency - webhooks, pagination, rate limits, and the patience to build robust clients around finicky or poorly documented APIs
  • Operational ownership - you instrument what you ship, notice when a pipeline degrades before anyone flags it, and fix root causes rather than restarting jobs
  • Independence - you can take a fuzzy operational problem, write the design doc, break it into tickets, and ship it end-to-end

Nice To Haves

  • Browser automation and scraping at scale - Playwright/Puppeteer, anti-fragile selectors, session management, structured crawling
  • Email infrastructure - deliverability, SPF/DKIM/DMARC, threading and reply-matching, inbox provider APIs
  • Building internal tools for operations teams - workflow software, queue-based work assignment, role-based access control, audit trails
  • Python and/or Go
  • Workflow orchestration - Airflow or similar
  • S3/object-storage-centric data architectures
  • OCR and document AI beyond LLM calls
  • Compliance-aware engineering - SOC 2, least-privilege access, action traceability
  • Experience with government, civic, or public-records data (Public records familiarity is a genuine plus)

Responsibilities

  • Own the systems behind NationGraph's unique differentiated asset: the data acquisition platform that collects public information/records at scale, (purchase orders, contracts, various procurement documents) from 100k+ government entities across the country.
  • Build high-volume interaction and ingestion pipelines with retries, dead-letter handling, checkpointing, and resumability.
  • Turn messy government documents (PDFs, spreadsheets, scans, portals) into clean structured data.
  • Build with AI at the core: LLM-driven document extraction, classification, and form processing at scale.
  • Design human-in-the-loop systems.
  • Contribute to the internal platform operations teams work in every day, complying with access controls, audit logging, and workflow tooling meeting SOC 2 requirements.
  • Own problems end to end in a small, senior team.
  • Take a goal like "automate records requests through web portals" from research to design doc to production.

Benefits

  • Competitive salary + early-stage equity
  • Unlimited PTO
  • High-quality health insurance, dental & vision coverage
  • Company provided lunches
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service