Senior Data Engineer

Node.DigitalWashington, DC
Remote

About The Position

We are seeking a Senior Data Engineer to provide authoritative expertise on data engineering methods and best practices, including code-first development approaches and modern pipeline design patterns. The role involves designing, implementing, and maintaining the data architecture that supports products and end users, with all assets managed under source control. Key responsibilities include developing ELT and ETL pipelines for efficient data processing in Azure Synapse and Azure Machine Learning, migrating source data, normalizing entity attributes, and reviewing/improving existing architecture. The position also requires establishing quality controls, incorporating source control, optimizing data processing and storage, developing self-service capabilities for analysts, authoring standard operating procedures, producing detailed documentation, and maintaining/expanding the environment. Staying current with emerging AI tooling for data engineering is also a part of this role.

Requirements

  • Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field. Alternatively, five years of applied work experience in any of the same fields.
  • 5 years - Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.
  • 5 years - Designing, implementing, and maintaining ELT and ETL processes in cloud based data analytics environments.
  • 3 years - Working in Azure Synapse and Azure Machine Learning with the modern data stack.
  • 3 years - Manipulating data in Python. Pandas is required.

Nice To Haves

  • DP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification.
  • Implementing pipelines and infrastructure using code first approaches: Python SDK, CLI, REST APIs, or infrastructure as code tooling such as Terraform or Bicep.
  • Implementing source control and continuous integration and delivery workflows for data assets.
  • Demonstrated familiarity with AI coding assistants and large language model integration patterns.
  • PySpark or Polars at production scale.
  • Entity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.
  • Building self service analytic access for non engineering users.
  • Certifications preferred, DP-203 or equivalent.
  • PySpark and Polars preferred.
  • Experience developing reusable, modular code preferred.

Responsibilities

  • Provide authoritative expertise on data engineering methods and best practices, including code first development approaches and modern pipeline design patterns.
  • Design, implement, and maintain the data architecture that supports products and end users, with all assets managed under source control.
  • Design, implement, and maintain ELT and ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning, using both SDK V1 and SDK V2.
  • Migrate source data identified by SBA OIG into Azure Data Lake Storage.
  • Normalize entity attributes such as addresses, phone numbers, and other common fields.
  • Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.
  • Establish quality controls across all pipelines and introduce error handling, logging mechanisms, and validation checks.
  • Incorporate source control across all pipelines and analytics codebases so code can evolve iteratively without destabilizing the architecture.
  • Optimize ingestion, processing, and storage across a wide variety of datasets and data types, including modern columnar formats such as Parquet.
  • Develop self service capabilities that let SBA OIG analysts query and export data for investigations and audits.
  • Author robust standard operating procedures governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.
  • Produce detailed documentation of the data architecture, including data dictionaries, entity relationship diagrams, and pipeline process maps.
  • Maintain and expand the environment with additional datasets and services on request, following a defined intake and testing process before production deployment.
  • Stay current with emerging AI tooling relevant to data engineering and contribute to exploratory work evaluating automation and language model assisted capabilities.

Benefits

  • Medical
  • Dental
  • Vision
  • Basic Life
  • Health Saving Account
  • 401K matching
  • Three weeks of PTO/Sick
  • 11 Paid Holidays
  • Pre-Approved Online Training
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service