Sr. Data Engineer

Knowledge Management, Inc.Washington, DC
Remote

About The Position

The Senior Data Engineer designs, builds, and maintains the data architecture and pipelines that support SBA OIG's Technology Solutions Division (TSD) in its loan fraud detection and investigative mission. The role works within SBA's Microsoft Azure cloud environment to migrate, transform, and structure data assets so that TSD's Data Analytics team can perform agile inquiries and systematic machine learning in support of audits and investigations.

Requirements

  • Five (5) years of hands-on experience maintaining SQL databases and conducting advanced operations in SQL and T-SQL.
  • Five (5) years of hands-on experience designing, implementing, and maintaining ELT/ETL processes in cloud-based data analytics environments.
  • Three (3) years of hands-on experience working in Azure Synapse and Azure Machine Learning within the modern data stack; DP-203 or an equivalent certification is preferred.
  • Three (3) years of hands-on experience manipulating data in Python, with required proficiency in Pandas; PySpark or Polars experience is preferred, along with experience developing reusable, modular code.
  • A Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field satisfies the education standard set forth in PWS Section 5.2.1.
  • In the absence of a bachelor's degree, five (5) years of applied work experience in data engineering, computer science, data science, machine learning, mathematics, or a related field will satisfy this requirement.
  • Microsoft Certified: Azure Data Engineer Associate, or an equivalent Azure data engineering certification, is preferred consistent with the Azure Synapse and Azure Machine Learning experience.

Nice To Haves

  • Implementing pipelines and infrastructure using code-first approaches, including Python SDK, CLI, REST APIs, or infrastructure-as-code tooling.
  • Implementing source control and CI/CD workflows.
  • Demonstrated familiarity with AI coding assistants and large language model integration patterns.
  • DP-203 or an equivalent certification is preferred.
  • PySpark or Polars experience is preferred.
  • Certifications supporting source control, CI/CD, or infrastructure-as-code practices are preferred but not required.

Responsibilities

  • Provide highly skilled and authoritative expertise on data engineering methods and best practices, including code-first development approaches and modern pipeline design patterns.
  • Design, implement, and maintain an efficient, secure, stable, and flexible data architecture, with all assets managed via source control.
  • Design, implement, and maintain ELT/ETL pipelines for processing source data in Azure Synapse and Azure Machine Learning.
  • Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.
  • Establish quality controls for pipeline maintenance, and introduce error handling, logging mechanisms, and validation checks.
  • Incorporate source control for all pipelines and data analytics codebases to enable iterative code development while maintaining data architecture stability.
  • Optimize the ingestion, processing, and storage of a wide variety of datasets and data types, including modern columnar formats such as Parquet.
  • Develop self-service capabilities for SBA OIG analysts to query and export data for investigations and audits.
  • Coordinate with the Senior Data Scientists to ensure the architecture supports machine learning algorithms and data pipelines in Azure Machine Learning.
  • Develop standard operating protocols governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.
  • Provide detailed documentation of the data architecture, including data dictionaries, entity-relationship diagrams, and pipeline process maps.
  • Maintain and expand the environment with additional datasets and services upon request, following a defined intake and testing process prior to production deployment.
  • Stay current with emerging AI tools relevant to data engineering and contribute to exploratory efforts evaluating automation and large language model-assisted capabilities.

Benefits

  • Health, dental, and vision insurance
  • 401(k) retirement plan
  • Paid time off (PTO) and holidays
  • Group Term Life and Accidental Death and Dismemberment Insurance
  • Voluntary Term Life Insurance
  • Short and Long-term disability insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service