Data Engineer (AWS, Spark)

OpenDataJobsWashington, DC
Hybrid

About The Position

At Peregrine Advisors, you will build the pipelines that move a federal agency's data from source to platform. As the work evolves, you will learn new systems and tools, take on greater responsibility, and help develop the firm's capabilities, tools, and lines of business. We move our best to where the hardest problems are. This is a full-time W-2 position with a salary of $103,000 to $140,000 per year. It is a hybrid work arrangement based in the Washington, DC metropolitan area, and the commuting cadence varies by assignment. The initial engagement requires United States citizenship and the ability to obtain a Public Trust determination. We are a data and technology innovation hub and a Benefit Corporation working at the center of the federal government's mission to deliver for client stakeholders and the US public.

Requirements

  • 4+ years of data engineering experience
  • Bachelor's degree
  • Spark ETL on AWS (Glue, Amazon EMR) in Python and PySpark
  • S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB
  • Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring
  • Data quality, validation, and lineage
  • Infrastructure-as-code (CloudFormation or Terraform)
  • Basic proficiency in writing, PowerPoint, and Excel

Nice To Haves

  • Master's degree in a relevant field
  • Trino or comparable federated SQL across the lake and relational stores
  • Apache Ranger-governed access
  • Legacy ETL migration (for example DataStage)
  • Federal information technology or high-volume data experience
  • Familiarity with AI-assisted developer tooling

Responsibilities

  • Build Spark-based extract, transform, and load (ETL) pipelines with Glue, Amazon EMR, Lambda, and Step Functions.
  • Write the processing in Python and PySpark.
  • Design the S3 layer, including Parquet, partitioning, and lifecycle, feeding Apache Iceberg tables.
  • Connect PostgreSQL on Amazon Aurora and DynamoDB, with Trino for federated Structured Query Language (SQL) across them.
  • Deliver pipelines that keep data clean and trustworthy at the volumes a federal platform runs at.
  • Automate ingestion that used to be handled case by case, and build the storage design everything downstream depends on.

Benefits

  • Medical, dental, and vision with the employee premium fully paid and half of dependent premiums
  • Employer-paid life, accidental death, and short-term and long-term disability insurance
  • A 401(k) matched 100% up to 4% of salary, vesting immediately
  • Unlimited paid time off
  • Sponsored professional certifications and continuing education
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service