Founding Engineer - Data

Mason•San Francisco, CA

About The Position

Mason is looking for a Founding Engineer - Data to join their team. This role is crucial in building the company's "brain" by transforming fragmented data from documents, databases, emails, and operational knowledge into connected data and context that AI agents can utilize. The engineer will design and build backend systems for ingesting, organizing, connecting, and retrieving enterprise knowledge, addressing challenges like inconsistent naming conventions, contradictory documents, and dispersed context across various formats. This position blends backend engineering, data architecture, and experimentation, working closely with the CTO and Infrastructure and Deployment engineers. The role involves owning production pipelines and processing infrastructure with a focus on data modeling, reconciliation, retrieval, and context quality.

Requirements

  • At least four years of engineering experience, with meaningful ownership of production backend systems, data platforms, or complex data integrations.
  • Strong programming and SQL skills, with depth in system design, databases, and data modeling.
  • A track record of building and operating pipelines that remain correct as sources, schemas, and volumes change.
  • Experience working with messy, heterogeneous data and resolving problems beyond the happy path.
  • Strong systems and algorithmic intuition, including tradeoffs around scale, correctness, incremental updates, and cost.
  • The ability to test an approach on real data, measure its limitations, and turn a promising experiment into a dependable system.
  • High ownership, high agency, and curiosity about how your data work improves the downstream product.

Nice To Haves

  • Experience with document processing, information retrieval, entity resolution, knowledge graphs, or LLM context systems is a strong plus.

Responsibilities

  • Design and build backend systems that ingest, organize, connect, and retrieve decades of enterprise knowledge.
  • Develop ingestion and processing pipelines for diverse data sources including databases, PDFs, policies, drawings, spreadsheets, and emails.
  • Create backend services and shared data infrastructure to transform customer integrations into reusable capabilities.
  • Design data models and entity relationships to connect information across fragmented systems.
  • Implement parsing, extraction, deduplication, and reconciliation methods for messy real-world data.
  • Build storage, indexing, and retrieval systems that provide agents with relevant context while preserving sources and access permissions.
  • Develop evaluation and observability systems to measure extraction quality, freshness, retrieval accuracy, latency, and cost.
  • Design processing architectures capable of scaling to billions of rows and large document collections with reliable updates and reprocessing.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service