Senior Software Engineer, Data Platform

Matterworks, Inc.Somerville, MA
Hybrid

About The Position

Matterworks is building foundation models for biochemical biology, akin to AlphaFold for proteins. As a Senior Software Engineer on the Data Platform team, you will be instrumental in constructing the connective tissue of our data platform. This involves designing, building, and scaling systems to transform raw samples and information into enriched, usable datasets with biological context. The data you and your team produce will power both our ML research, model development, and product offerings. You will report to the Head of Engineering and collaborate daily with machine learning researchers, scientists, and the product team.

Requirements

  • Significant professional experience building production data systems and pipelines (scope and judgment are prioritized over years of experience).
  • Proficient in Python and SQL for large-scale data processing.
  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Experience with Airflow or Dagster is transferable.
  • Demonstrated experience designing stable identifiers for a large, changing corpus and building validation that gates a publish rather than reporting on it after the fact.
  • Experience integrating an LLM or agent component into a production data path, including the evaluation loop, gold set, and cost per record.
  • Daily use of AI coding tools, coupled with a healthy skepticism regarding their output on production data correctness.
  • A track record of owning work through to a running, validated system, including addressing issues like malformed datasets, writing backfills, and debugging failed processes.
  • Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar). Strong engineering depth is required.
  • A passion for contributing to an early-stage startup, valuing autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges over rigid processes and egos.

Responsibilities

  • Build and Scale Data Contracts: Own the pipelines and systems that other teams consume. Design and implement systems that effectively scale to multiple petabytes of data.
  • Serving and Cost at Scale: Build and scale systems for acquiring, storing, and serving data efficiently and affordably as it grows, focusing on data layout, featurization throughput, Kubernetes-native orchestration, and cost management.
  • Labels and Enrichment: Transform raw data into usable datasets with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features.
  • Quality Gates: Automate quality checks to enable increasing capability without regression, promoting only on successful completion.
  • Interfaces People Use: Own the interfaces that AI, chemistry, product, and agents interact with, from the SDK for building datasets to the tools exposing platform capabilities.
  • Operations and Data Rights: Ensure the effective operation of our data needs, meeting designed service level agreements while providing high-quality, provenance, and secure data processing in line with customer needs.

Benefits

  • Competitive base salary
  • Stock options
  • Health & dental insurance
  • Vision insurance
  • Long- and short-term disability insurance
  • Life insurance
  • 401k with company match
  • Flexible work policy
  • Unlimited time away policy
  • Commuter benefits
  • Parking
  • Regular team meals and outings
  • Company support for continued education/coursework
  • Company support for conference participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service