Data Engineer

Edgesource CorporationMcLean, VA
$130,000 - $145,000Hybrid

About The Position

We are seeking a Data Engineer to work with a small team to build complex data flows for a custom application. Successful candidate will have advanced Python programming skills, familiarity with Java, an understanding of data security, privacy, governance and compliance principles and a demonstrated history of building production data pipelines and ETL workflows at scale.

Requirements

  • Active TS/SCI Full Scope Polygraph
  • Minimum of 5 years' experience
  • Demonstrated experience building production data pipelines and ETL/EL workflows at scale
  • Proficiency with Apache Spark and PySpark for distributed data processing
  • Advanced Python programming skills including data manipulation libraries (Pandas, NumPy) and data engineering best practices
  • Understanding of data security, privacy, governance, and compliance principles
  • Experience with workflow orchestration tools (such as Step Functions, Airflow)
  • Familiarity with containerization (such as Docker or Podman) and deploying data applications in cloud environments
  • Experience with AWS services (S3, Lambda, Step Functions)
  • Experience with PostgreSQL and MySQL in production environments, including performance tuning and schema design
  • Demonstrated experience with SQL and query optimization for complex analytical workloads
  • Experience with version control (Git) and Cl/CD practices for data pipelines
  • Demonstrated ability to work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
  • Strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks

Nice To Haves

  • Experience with data Lakehouse architecture using Apache Iceberg
  • Hands-on experience configuring, deploying, and integrating data platform components: Apache Ranger (access control and data governance), Trino (distributed SQL query engine), Data catalogs (Unity Catalog OSS, Apache Polaris, etc.), Apache Superset data visualization and dashboarding)
  • Proficiency with Bash scripting for automation and data processing tasks
  • Experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure
  • Familiarity or experience with tracking data lineage and associated tooling such as Open lineage
  • Familiarity or experience with Java
  • Familiarity with data quality frameworks, testing methodologies, and validation strategies
  • Background with large-scale data migrations or platform modernization efforts
  • Experience integrating Al/ML services and models (translation, OCR, speech-to-text, NLP, language detection, topic modeling), LLMs, and RAG retrieval-augmented generation) pipelines
  • Familiarity with geospatial data processing H3, PostGIS, or similar)
  • Contributions to data engineering documentation, best practices, and design patterns
  • Experience with NoSQL databases (DynamoDB, etc.)

Responsibilities

  • Building end-to-end data pipelines leveraging Python
  • Using orchestration tools to deploy data pipelines, including configuring and updating Spark Jobs
  • Containerizing and deploying applications in cloud environments like AWS.
  • Working with MySQL and PostgreSQL including performance tuning, schema design, and query optimization for complex, analytical workloads.
  • Leveraging industry standard tools for code control (Git, IaaC control, etc.)
  • Working with data catalogs, tracking data lineage and handling a variety of data formats, including Geospatial.
  • Using Bash scripting for automation and data processing tasks
  • Integrating Al/ML services and models
  • Work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
  • Leverage strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks
  • Leverage a background in large-scale data migration or platform modernization efforts
  • Contribute to data engineering documentation, best practices, and design patterns.

Benefits

  • Flexible PTO Policy + 11 Paid Holidays
  • Flexible Work Schedules (Remote / Hybrid)
  • Medical / Dental / Vision / Flexible Spending Account (FSA)
  • 401k Plan with Match
  • Tuition & Professional Development Support
  • Commuter Benefits
  • Bonus & Employee Referral Programs
  • Career Growth Opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service