Data Engineer Lead

CapgeminiAtlanta, GA

About The Position

We are seeking a Data Engineer with strong expertise in Google Cloud Platform (GCP) and PySpark to design, build, and optimize scalable data pipelines. The role focuses on processing large-scale datasets, enabling analytics, and supporting data-driven decision-making within a cloud-native ecosystem. The ideal candidate will have hands-on experience with distributed data processing, cloud data services, and ETL/ELT frameworks, with a strong engineering mindset toward performance, scalability, and reliability.

Requirements

  • Strong Python + PySpark
  • Experience with Spark (RDD/DataFrame APIs)
  • Hands-on Google Cloud Platform (GCP): BigQuery, Cloud Storage, Dataproc / Dataflow
  • ETL/ELT pipeline development
  • Data modeling (structured & semi-structured data)
  • SQL (advanced)
  • Airflow / Composer (workflow orchestration)
  • CI/CD pipelines
  • Strong data engineering mindset (not just scripting)
  • Experience with large-scale datasets (TB-level)
  • Comfortable working in cloud-native, distributed environments
  • Proactive in debugging and optimizing data pipelines

Nice To Haves

  • Pub/Sub (preferred)
  • Streaming frameworks (Kafka / Pub-Sub streaming)
  • Delta Lake / Iceberg / Lakehouse patterns
  • Machine Learning pipeline exposure
  • Data governance and lineage tools

Responsibilities

  • Design and develop scalable data pipelines using PySpark
  • Process and transform large datasets (batch and streaming)
  • Build reusable data processing frameworks
  • Work extensively with GCP services, including: BigQuery (data warehousing), Cloud Storage (data lake), Dataflow / Dataproc (processing)
  • Optimize ingestion, storage, and retrieval of datasets in GCP
  • Develop and maintain end-to-end ETL/ELT pipelines
  • Ensure data quality, data consistency, and schema evolution handling
  • Optimize PySpark jobs for large-scale distributed processing, memory and execution efficiency
  • Tune query performance in BigQuery
  • Collaborate with data scientists, analysts, and application teams
  • Enable datasets for analytics, reporting, and ML workflows
  • Implement CI/CD for data pipelines
  • Use Git for version control
  • Automate deployments and monitoring of data workflows

Benefits

  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade
  • Company paid holidays
  • Personal Days
  • Sick Leave
  • Medical, dental, and vision coverage
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service