Data Engineer Role

OpenDataJobsWashington, DC

About The Position

The work Data Engineers build and operate the systems that move data from its sources to the people, applications, analyses, and models that depend on it. They ingest data, transform it, organize it for use, and keep it accurate, secure, traceable, and available at scale. Their work turns fragmented files, documents, databases, application programming interfaces (APIs), and event streams into reliable data products. Artificial intelligence (AI) is one important consumer of that work, alongside reporting, visualization, analytics, software, and operational systems. Some openings may involve preparing dependable data for machine learning, document retrieval, or AI evaluation. The center of the Role remains dependable data engineering from source to use. What you may build: Ingestion and transformation pipelines for batch, streaming, and event-driven data from APIs, databases, files, documents, object stores, messaging systems, and operational platforms. Cloud and on-premises data platforms, including databases, data lakes, warehouses, lakehouses, and serving layers for reporting, visualization, software, analytics, and other operational uses. Quality, validation, metadata, lineage, provenance, and access-control capabilities that make data trustworthy, explain how it changed, and keep its use within approved boundaries. The operational layer around data products: orchestration, testing, monitoring, backfills, replay, recovery, retention and deletion implementation, performance and cost tuning, infrastructure as code, and technical documentation. Where the work requires it, versioned feature, training, testing, or evaluation datasets; document and retrieval-index pipelines; or governed telemetry and feedback data that support machine learning and generative AI systems.

Requirements

  • Care about data completeness, timeliness, understanding, authorization, and fitness for use.
  • Trace failures across sources, transformations, storage, and serving layers.
  • Improve recurring processes instead of working around them.
  • Collaborate well with source-system owners, software engineers, analysts, data scientists, AI Engineers, Machine Learning Engineers, security and governance specialists, and DevSecOps Engineers.
  • Make data contracts and tradeoffs clear.
  • Distinguish a data problem from a model or application problem.
  • Prefer ownership of an outcome to a narrowly assigned task.
  • A working foundation in programming and query languages used for data engineering (Python and SQL are common).
  • Experience or strong grounding in data ingestion, transformation, storage, schema and data-model design, and the performance characteristics of distributed data systems.
  • An understanding of batch, streaming, and event-driven processing, together with orchestration, testing, deployment, monitoring, recovery, and documentation.
  • Practical experience with data quality, metadata, lineage, provenance, versioning, access controls, and secure data handling.
  • Judgment to work in environments where accuracy, privacy, security, traceability, reproducibility, resilience, performance, and cost matter.

Responsibilities

  • Build and operate systems that move data from its sources to consumers.
  • Ingest, transform, and organize data for use.
  • Ensure data is accurate, secure, traceable, and available at scale.
  • Turn fragmented files, documents, databases, APIs, and event streams into reliable data products.
  • Prepare dependable data for machine learning, document retrieval, or AI evaluation.
  • Build ingestion and transformation pipelines for batch, streaming, and event-driven data.
  • Develop cloud and on-premises data platforms (databases, data lakes, warehouses, lakehouses, serving layers).
  • Implement quality, validation, metadata, lineage, provenance, and access-control capabilities.
  • Manage the operational layer around data products (orchestration, testing, monitoring, backfills, replay, recovery, retention, deletion, performance tuning, infrastructure as code, documentation).
  • Create versioned datasets, document and retrieval-index pipelines, or governed telemetry and feedback data where required.

Benefits

  • Compensation, benefits, work location, and employment terms are set for each specific opening and will be stated with that opening.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service