Data Engineer Junior

ASM ResearchSan Antonio, TX

About The Position

The Junior Data Engineer supports the design, development, and maintenance of data pipelines and infrastructure that enable efficient data ingestion, processing, and storage within a regulated enterprise environment. The role focuses on building scalable data workflows, implementing transformation logic, and ensuring data quality and integrity across mission‑critical systems for analytics and reporting. The engineer collaborates with analysts, data scientists, and security teams to deliver reliable datasets, contribute to data governance and lineage tracking, and drive continuous improvement of data platform performance and reliability.

Requirements

  • Bachelor’s degree in Computer Science, Information Systems, Data Engineering, or a related field; or equivalent relevant experience in data engineering or data development.
  • Typically 1–3 years of experience in data engineering, data development, or a closely related field supporting data pipelines, databases, or analytics platforms.
  • Hands‑on experience with Python and SQL for building data pipelines, transformations, and queries; exposure to distributed processing frameworks such as Spark is strongly preferred.
  • Experience working with both relational and non‑relational data stores (e.g., SQL databases, NoSQL, data lakes, or warehouses) in an enterprise or federal IT environment.
  • Familiarity with data quality practices, including validation rules, automated testing, and monitoring of data reliability.
  • Ability to obtain and maintain a SECRET‑level security clearance as required by the client environment.
  • U.S. Citizenship required to meet federal staffing and clearance requirements.

Nice To Haves

  • Experience with cloud data services on platforms such as AWS, Azure, or Google Cloud (e.g., managed databases, data lakes, orchestration services).
  • Familiarity with big data ecosystems such as Hadoop or Spark, including experience tuning jobs or managing distributed data processing.
  • Exposure to data governance frameworks, metadata management, and lineage tools in regulated or federal environments.
  • Experience working with workflow orchestration tools such as Apache Airflow or similar schedulers in production environments.

Responsibilities

  • Develop, operate, and maintain automated ETL/ELT pipelines using tools such as Python, SQL, or Spark to process structured and unstructured data across enterprise data platforms.
  • Implement data models, schemas, and storage solutions in relational and non‑relational databases, including data lakes and data warehouses, to support analytics and reporting requirements.
  • Apply data quality validation techniques, including anomaly detection, schema enforcement, and automated testing frameworks, to ensure accuracy, completeness, and consistency of data.
  • Support orchestration and scheduling of data pipelines using workflow managers such as Apache Airflow or similar tools, ensuring reliable, repeatable execution.
  • Optimize data processing performance through indexing, partitioning, and distributed computing approaches to meet latency and throughput targets.
  • Implement data security controls, including encryption, access management, and adherence to applicable federal data protection standards, in coordination with security teams.
  • Track and document data lineage and metadata in support of data governance, auditability, and regulatory reporting expectations.
  • Troubleshoot pipeline failures, latency issues, and data inconsistencies across integrated systems, coordinating with cross‑functional teams to resolve root causes.
  • Collaborate with data scientists, analysts, and business stakeholders to translate analytics requirements into scalable, maintainable data solutions and improvements to the data platform.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service