Data Engineer – Baltimore City, MD

Creative Information Technology, Inc.Falls Church, VA
Onsite

About The Position

Creative Information Technology Inc (CITI) is an esteemed IT enterprise renowned for its exceptional customer service and innovation. We serve both government and commercial sectors, offering a range of solutions such as Healthcare IT, Human Services, Identity Credentialing, Cloud Computing, and Big Data Analytics. With clients in the US and abroad, we hold key contract vehicles including GSA IT Schedule 70, NIH CIO-SP3, GSA Alliant, and DHS-Eagle II. Join us in driving growth and seizing new business opportunities. Client is seeking a hands-on Data Engineer to design, develop, and optimize large-scale data pipelines in support of our Enterprise Data Warehouse (EDW) and Data Lake solutions. This role requires deep technical expertise in coding, pipeline orchestration, and cloud-native data engineering on AWS. The Data Engineer will be directly responsible for implementing ingestion, transformation, and integration workflows — ensuring data is high-quality, compliant, and analytics-ready. This role may support other projects or teams within MDH as needed.

Requirements

  • Experience as a data engineer or similar role with a strong understanding of data architecture and ETL processes.
  • Proficient in programming languages for data processing.
  • Knowledgeable of distributed computing and parallel processing.
  • 3+ years hands-on experience in building, deploying, and maintaining data pipelines on AWS or equivalent cloud platforms.
  • Strong coding skills in Python and SQL (Scala or Java a plus).
  • Proven experience with Apache Spark (PySpark) for large-scale processing.
  • Hands-on experience with AWS Glue, S3, Redshift, Athena, EMR, Lake Formation.
  • Strong debugging and performance optimization skills in distributed systems.
  • Hands-on experience with Iceberg, Delta Lake, or other OTF table formats.
  • Experience with Airflow or other pipeline orchestration frameworks.
  • Practical experience in CI/CD and Infrastructure-as-Code (Terraform, CloudFormation).
  • Practical experience with EDI X12, HL7, or FHIR data formats.
  • Strong understanding of Medallion Architecture for data lake houses.
  • Hands-on experience building dimensional models and data warehouses.
  • Working knowledge of HIPAA and CMS interoperability requirements.

Nice To Haves

  • Scala or Java a plus.

Responsibilities

  • Designing, building, and maintaining data pipelines and infrastructure to support data-driven decisions and analytics.
  • Designing, developing and maintaining data pipelines, and extract, transform, load (ETL) processes to collect, process and store structured and unstructured data.
  • Building data architecture and storage solutions, including data lakehouses, data lakes, data warehouse, and data marts to support analytics and reporting.
  • Developing data reliability, efficiency, and quality checks and processes.
  • Preparing data for data modeling.
  • Monitoring and optimizing data architecture and data processing systems.
  • Collaboration with multiple teams to understand requirements and objectives.
  • Administering testing and troubleshooting related to performance, reliability, and scalability.
  • Creating and updating documentation.
  • Designing, coding, and deploying ETL/ELT pipelines across bronze, silver, and gold layers of the Data Lakehouse.
  • Building ingestion pipelines for structured (SQL), semi-structured (JSON, XML), and unstructured data using PySpark/Python programming language using AWS Glue or EMR.
  • Implementing incremental loads, deduplication, error handling, and data validation.
  • Actively troubleshooting, debugging, and optimizing pipelines for scalability and cost efficiency.
  • Developing dimensional data models (Star Schema, Snowflake Schema) for analytics and reporting.
  • Building and maintaining tables in Iceberg, Delta Lake, or equivalent OTF formats.
  • Optimizing partitioning, indexing, and metadata for fast query performance.
  • Building ingestion and transformation pipelines for EDI X12 transactions (837, 835, 278, etc.).
  • Implementing mapping and transformation of EDI data with FHIR and HL7 frameworks.
  • Working hands-on with AWS Health Lake (or equivalent) to store and query healthcare data.
  • Developing automated validation scripts to enforce data quality and integrity.
  • Implementing IAM roles, encryption, and auditing to meet HIPAA and CMS compliance standards.
  • Maintaining lineage and governance documentation for all pipelines.
  • Working closely with the Lead Data Engineer, analysts, and data scientists to deliver pipelines that support enterprise-wide analytics.
  • Actively contributing to CI/CD pipelines, Infrastructure-as-Code (IaC), and automation.
  • Continuously improving pipelines and adopting new technologies where appropriate.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service