Research Engineer, Data Infrastructure (Language Modeling)

Cartesia•San Francisco, CA
•$200,000 - $350,000•Onsite

About The Position

Data is the lifeblood of our models, and we are looking for a Research Engineer, Data Infrastructure to build the datasets and systems that power pretraining at Cartesia. In this role, you will write performant, scalable infrastructure to acquire, process, and curate massive datasets, and partner closely with research to optimize the characteristics and composition of data mixtures. Your work will directly shape the capabilities and quality of our foundational models.

Requirements

  • Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training and inference.
  • Strong modern engineering execution: clean, well-tested code, fluency with current tools, and a willingness to pick the right tool for the problem rather than defaulting to familiar patterns.
  • Familiarity with building and evaluating datasets for generative models and reasonable working knowledge of how they're trained and inference.

Nice To Haves

  • Experience with large-scale data processing using parallel infrastructure such as Ray, Spark, or Kubernetes.
  • Experience with pretraining language models.

Responsibilities

  • Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets.
  • Design and operate scalable, high-throughput, and reproducible data pipelines — covering ingestion, preprocessing, filtering, deduplication, and augmentation.
  • Design and run ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality.
  • Partner closely with research and infrastructure teams to co-design data loading, versioning, and experimentation pipelines.
  • Establish and enforce rigorous standards for data quality, with a tight feedback loop between dataset characteristics and model behavior.
  • Identify and source novel datasets; manage relationships and budgets with external data vendors and partners.

Benefits

  • Competitive base salary alongside attractive equity package.
  • Fully covered medical insurance along with dental and vision for you and your family.
  • 9 weeks paternity & 12 weeks maternity leave
  • 401(k)
  • A monthly stipend to help you get to and from the office.
  • Flexible PTO Take as much time as you need to recharge your batteries.
  • Lunch, dinner and plenty of snacks, provided daily.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service