Senior Software Engineer, Data Platform

MatterworksSomerville, MA
Hybrid

About The Position

Matterworks is building foundation models that make the 'dark matter' of biochemical biology legible. As a Senior Software Engineer, you will build the connective tissue of our data platform, designing, building, and scaling systems to enrich data from raw samples and information into readily usable datasets enriched with biological context. The data produced will serve our ML research, model development, and product. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.

Requirements

  • Significant professional experience building production data systems and pipelines.
  • Proficient in Python and SQL for large-scale data processing.
  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Airflow or Dagster experience transfers fine.
  • Demonstrated experience designing stable identifiers for a large, changing corpus, and building validation that gates a publish rather than reporting on it after the fact.
  • Experience putting an LLM or agent component into a production data path, including the eval loop, the gold set, and cost per record.
  • Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.
  • A track record of owning work through to a running, validated system, including fixing malformed datasets, writing backfills, and debugging bad publishes.
  • Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar).
  • Engineering depth is the requirement.
  • A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.

Responsibilities

  • Build and Scale Data Contracts: Own the pipelines and systems other teams consume from. Design and implement systems that scale to multiple petabytes of data effectively.
  • Serving and Cost at Scale: Build and scale systems to acquire, store, and serve data quickly and affordably as it grows, focusing on data layout, featurization throughput, Kubernetes-native orchestration, and cost management.
  • Labels and Enrichment: Turn raw data into usable datasets with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features.
  • Quality Gates: Automate quality checks to enable increasing capability without regression, promoting only on a pass.
  • Interfaces People Use: Own the surfaces AI, chemistry, product, and agents call, from the SDK used to build datasets to the tools that expose platform capabilities.
  • Operations and Data Rights: Ensure effective operations of data needs, meeting designed service level agreements, while providing high quality, provenance, and secure data processing in line with customer needs.

Benefits

  • Competitive base salary
  • Stock options
  • Health & dental insurance
  • Vision insurance
  • Long- and short-term disability insurance
  • Life insurance
  • 401k with company match
  • Flexible work policy
  • Unlimited time away policy
  • Commuter benefits
  • Parking
  • Regular team meals and outings
  • Company support for continued education/coursework
  • Company support for conference participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service