Research Data Engineer

DRCNew York, NY
$95,000 - $130,000Remote

About The Position

We're looking for someone to help us build a unified system for our primary research data, financial data, and business development information — including reconciling older historical data with newer data that has drifted in format or structure over time, so we can run reliable long-term analyses. This role is a strong fit for someone with a background in library or information science combined with programming experience. A large part of this job is about organizing, documenting, and making sense of data that has changed shape over the years — the kind of problem information science trains people to solve well. You'll also build the technical pipelines and API connections that move this data between our systems, with support available for the more infrastructure-heavy pieces (cloud setup, deployment) as needed. This is a backend-focused role — no custom dashboard or front-end UI work required. We use off-the-shelf BI tools for reporting.

Requirements

  • A degree or coursework in Library Science, Information Science, or a related field (MLIS or equivalent experience)
  • Demonstrated programming experience — coursework, personal projects, prior job, or research work involving Python and/or SQL
  • Strong understanding of metadata, taxonomy, data provenance, and information organization principles
  • Comfort learning new technical tools and systems independently
  • Excellent documentation and communication skills — this role sits between technical and non-technical colleagues
  • Bachelor's degree required; graduate degree in library/information science strongly preferred

Nice To Haves

  • Experience with research data management, digital scholarship, or data curation in an academic or research library setting
  • Familiarity with cloud data platforms (Snowflake, BigQuery, AWS, or similar)
  • Exposure to ETL/orchestration tools (Airflow, dbt) or API development
  • Experience with survey, longitudinal, or panel data specifically

Responsibilities

  • Design how our historical and current data fit together, including resolving inconsistencies from changed formats, renamed fields, or different collection methods over time
  • Build a clear, well-documented data dictionary and provenance record — where data came from, how it's changed, and what it means today
  • Build data pipelines (ETL/ELT) to bring data in from internal tools and external vendors on a reliable schedule
  • Build and maintain API connections between our systems and third-party software (CRM, accounting/finance tools, research platforms, and similar)
  • Apply sensible data governance and access controls for sensitive financial and confidential records
  • Monitor data quality over time and flag issues before they affect analysis
  • Partner with IT/outside contractors as needed on cloud infrastructure and deployment, especially early on

Benefits

  • medical, dental, and vision insurance
  • paid time off
  • participation in the Firm’s 401(k) plan
  • additional employee benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service