Data Engineer

Princeton UniversityPrinceton, NJ
$120,000 - $135,000Onsite

About The Position

The Princeton DMIA Integration team is looking for a Data Engineer to own and evolve our data integration practice. You will be responsible for ingesting data from enterprise source systems into our data warehouse platform — what the CIO office refers to as system-to-data-repository integrations. You will play a key role in our migration to a cloud-native data platform, with candidates expected to bring expertise in modern tooling such as Microsoft Fabric, Snowflake with dbt and Fivetran, or Databricks. The existing data warehouse technologies include IBM DataStage, SQL, Oracle, and Shell-based pipelines running on on-premises Linux infrastructure.

Requirements

  • Linux / Shell scripting
  • Python programming
  • SQL (query authoring and data validation)
  • Basic networking (DNS, HTTP/S, TCP/IP, proxies, firewalls)
  • Cloud infrastructure fundamentals (any major provider)
  • IAM: LDAP, Active Directory, OAuth 2.0, certificate management
  • 5+ years of proven experience with enterprise ETL/ELT tooling
  • Advanced SQL skills across multiple platforms (Oracle, SQL Server, PostgreSQL)
  • Strong Python programming skills for data transformation, scripting, and pipeline automation
  • Shell scripting proficiency for batch job automation on Linux/Unix servers
  • Hands-on experience with at least one cloud data platform: Microsoft Fabric, Snowflake (with dbt and/or Fivetran), or Databricks
  • Experience with dbt (data build tool) for transformation layer development
  • Linux/Unix systems fluency including file management, cron scheduling, and process monitoring
  • Understanding of data warehousing concepts: dimensional modelling, star/snowflake schemas, slowly changing dimensions
  • Familiarity with basic networking, storage, and cloud infrastructure concepts
  • Experience with IAM and access control: LDAP, Active Directory, and database-level permission management
  • Working knowledge of REST APIs for source system data extraction and pipeline orchestration
  • Ability to document data flows, pipeline architecture, and transformation logic clearly
  • Bachelor’s degree in computer science

Nice To Haves

  • 7+ years proven experience with enterprise ETL/ELT tooling
  • Familiarity with orchestration tools such as Apache Airflow, Azure Data Factory, or Prefect
  • Exposure to streaming or near-real-time ingestion patterns (Kafka, Kinesis, Event Hubs)
  • Experience with data quality frameworks (Great Expectations, Soda, or equivalent)
  • Cloud platform certifications (Azure, AWS, or GCP data engineering tracks)

Responsibilities

  • Design, build, and maintain ETL/ELT pipelines to ingest, transform, and load data from source systems into the enterprise data warehouse.
  • Lead the migration of on-premises data pipelines to the organization’s future cloud-native data platform (one of Fabric, Snowflake + dbt + Fivetran, or Databricks).
  • Implement and enforce data quality checks, data lineage tracking, and pipeline observability across all integration workflows.
  • Ensure data security and compliance requirements are met, including encryption at rest and in transit, and access controls aligned with IAM policies.
  • Optimize pipeline performance, scheduling, and resource utilization across batch and incremental load patterns.
  • Partner with data analysts, BI developers, and source system owners to understand data requirements and translate them into robust ingestion pipelines.
  • Operate and support ingestion and transformation pipelines.
  • Develop and maintain data pipeline documentation, data dictionaries, and SLA agreements for ingestion jobs.
  • Participate in on-call production support rotation and respond to integration incidents per SLA.
  • Contribute to CI/CD pipeline setup and DevOps practices for data integration deployments.

Benefits

  • Comprehensive benefit program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service