Data/ML Engineer

Princeton University•Princeton, NJ
•$120,000 - $135,000•Remote

About The Position

The Accelerator seeks a Data/ML Engineer to strengthen its data team and advance the engineering, enrichment, and provisioning of the data it collects. The Accelerator at Princeton includes a portfolio of multiple planned independent and intersecting tools, built on a shared data and compute platform serving computational social scientists at research institutions across North America, Europe, and Africa. The Data/ML Engineer will work within the team to help drive data engineering and machine learning initiatives and collaborations. They will play a crucial role in building and operating the pipelines that transform large-scale social media and web behavior data into research-ready data products, and in developing the machine learning and enrichment capabilities that extend their value. They will work on problems that have no precedent and little source material, requiring novel solutions. They will also be responsible for working with the other teams within the Accelerator and external partners to help foster collaboration and create an impactful environment for users. This is a 6-month term role with potential for extension. A remote work arrangement within the United States may be considered for candidates with the appropriate background and experience.

Requirements

  • 3+ years of relevant experience as a data engineer, machine learning engineer, or data scientist, which may include graduate research and internship experience, with a record of building production systems that operate reliably at scale.
  • Experience working in a remote, agile environment.
  • Bachelor's degree or equivalent in a relevant field.
  • Strong proficiency in Python and SQL, and hands-on experience with distributed data processing (e.g., Apache Spark) on large data volumes.
  • Experience building, evaluating, and operating machine learning or NLP pipelines, including batch inference.
  • Working knowledge of cloud data platforms.
  • Strong communication and interpersonal skills to effectively collaborate with researchers in the field, other engineers at various levels of experience, and administrative and leadership team members.

Nice To Haves

  • Experience with Azure and Databricks, including Unity Catalog.
  • Experience with infrastructure-as-code (e.g., Terraform), containers, and CI/CD tooling.
  • Experience with large-scale social media, web behavior, or text-as-data research.
  • Familiarity with large language model annotation workflows and their evaluation.
  • Publications in reputable scientific journals or conferences is desirable.

Responsibilities

  • Work closely with the Accelerator leadership team to align data engineering and machine learning initiatives with overarching goals and long-term vision.
  • Identify and prioritize development projects that benefit from data engineering and machine learning methodologies and innovations.
  • Design, build, and operate data pipelines across the Accelerator's medallion architecture, with end-to-end ownership of transformation layers that serve researchers.
  • Ensure the accuracy, integrity, and quality of data to be made available through the Accelerator, including data quality validation at pipeline boundaries and enforcement of versioned schema contracts.
  • Diagnose and optimize distributed data processing workloads at production scale.
  • Develop deployment automation, CI/CD, and release processes for data products, including versioned data releases and researcher-facing change documentation.
  • Design, develop, and operate ML and NLP enrichment pipelines over large-scale text and behavioral data, including language identification, translation, and topic and content classification.
  • Own the full lifecycle of enrichment models: selection, evaluation against labeled data, batch inference architecture, cost efficiency, and reprocessing and versioning strategy.
  • Develop ML-ready feature layers and data products to support advanced research use cases.
  • Evaluate and apply large language model workflows and other emerging AI methods where they demonstrably improve outcomes, with attention to their validity for downstream scientific analysis.
  • Apply statistical analysis and modeling to characterize datasets, estimate coverage, and support research design.
  • Contribute to cost attribution, visibility, and governance across institutional workspaces, including cluster policies, budget controls, and storage lifecycle management.
  • Design data and ML workloads to operate within the platform's cost governance framework.
  • Develop automation for workspace and project provisioning as institutions and research projects onboard.
  • Operate within Unity Catalog governance, multi-tenant isolation, and research data security requirements.
  • Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
  • Author and maintain researcher-facing documentation and provide direct technical support to research users of the platform.
  • Collaborate with research teams to define data products, sampling frames, and enrichment requirements, and apply state-of-the-art techniques to ongoing scientific challenges.
  • Stay current with the latest advancements in data engineering, machine learning, and relevant fields to continuously innovate.
  • Build strong relationships with external partners, driving collaborations that enhance the Accelerator's scientific impact.

Benefits

  • Comprehensive benefit program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service