Data Engineer (AWS, Spark)

OpenDataJobsWashington, DC
Hybrid

About The Position

Peregrine Advisors is a firm founded on a simple conviction: the best solutions come from the people closest to the problem, given real ownership and the tools to deliver. We are a data and technology innovation hub and a Benefit Corporation working at the center of the federal government's mission to deliver for client stakeholders and the US public, looking for highly motivated contributors who thrive when trusted to own a hard problem and equipped to deliver the solution. Your first assignment is to build the pipelines that move a federal agency's data from source to platform on Amazon Web Services (AWS): ingest, process, store, and keep it clean and trustworthy at scale. The work is real, hard, and it matters. It is also where you start, not the shape of your career here: we hire people, not seats, and we move our best to where the hardest problems are.

Requirements

  • Sole United States citizenship and the ability to obtain a Public Trust determination are required for this initial engagement.
  • Hybrid role based in the Washington, DC metropolitan area, and it requires commuting into DC regularly.
  • 4+ years of data engineering experience
  • Bachelor's degree
  • Spark ETL on AWS (Glue, Amazon EMR) in Python and PySpark
  • S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB
  • Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring
  • Data quality, validation, and lineage
  • Infrastructure-as-code (CloudFormation or Terraform)
  • Basic proficiency in writing, PowerPoint, and Excel

Nice To Haves

  • Master's degree in a relevant field
  • Trino or comparable federated SQL across the lake and relational stores
  • Apache Ranger-governed access
  • Legacy ETL migration (for example DataStage)
  • Federal information technology or high-volume data experience
  • Familiarity with AI-assisted developer tooling

Responsibilities

  • Build ingest-process-store pipelines on AWS: Spark-based extract, transform, and load (ETL) with Glue, Amazon EMR, Lambda, and Step Functions, in Python and PySpark.
  • Develop a data lake that feeds the platform: S3 design (Parquet, partitioning, lifecycle) into Apache Iceberg tables, PostgreSQL on Amazon Aurora, and DynamoDB, with Trino for federated SQL across them, plus event orchestration, secrets and monitoring, and data quality, validation, and lineage built in.
  • Contribute to the firm itself in time: new capabilities, tools, and lines of business you help spin up.

Benefits

  • Medical, dental, and vision with the employee premium fully paid and half of dependent premiums
  • Employer-paid life, accidental death, and short-term and long-term disability insurance
  • A 401(k) matched 100% up to 4% of salary, vesting immediately
  • Unlimited paid time off
  • Sponsored professional certifications and continuing education
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service