Big Data Lead

HEXAWAREUnited States,

About The Position

This role requires a highly experienced Big Data Lead with advanced proficiency in Python and SQL, as well as expertise in AWS services. The ideal candidate will be responsible for developing and maintaining data pipelines, optimizing SQL queries, implementing testing and documentation, and managing version control. This is a hands-on role requiring strong communication and collaboration skills.

Requirements

  • 10+ Years Experience
  • Great Communicator/Client Facing/ Attention to detail
  • Individual Contributor and ability to work as a team.
  • 100% Hands on in the mentioned skills
  • Advanced Proficiency in Python concepts like Code Structures, Modules, Packages, Class, SubClass, Inheritance, Multi-Threading and Functional Programming.
  • Experience in developing reusable Python packages for internal or public usage
  • Ability to write automating ETL processes and scheduling jobs like Airflow DAG.
  • Ability to track job pipeline runs to reprocess error records
  • Ability to orchestration different pipelines to run in sequence or parallel.
  • Troubleshoot data pipeline errors and fix issues
  • Export or Import data to/from various formats like CSV, JSON, XML etc preferably from S3 or other cloud storage.
  • Experience in using AI IDE tool
  • Advanced SQL skills, including complex joins, CTE's and subqueries
  • Experience in optimizing SQL queries for performance and optimization in data warehouse technologies preferably Snowflake
  • Proficiency in Python unit, integration and system test.
  • Proficiency in implementing DBT tests for data validation and quality checks
  • Experience in generating code using configurations using python and jinja templates
  • Experience in GitHub, including implementing CI/CD process from scratch
  • In depth understanding of AWS S3 for data storage, ECS, IAM including best practices for organization and security.

Responsibilities

  • Develop reusable Python packages for internal or public usage.
  • Write automating ETL processes and scheduling jobs like Airflow DAG.
  • Track job pipeline runs to reprocess error records.
  • Orchestrate different pipelines to run in sequence or parallel.
  • Troubleshoot data pipeline errors and fix issues.
  • Export or import data to/from various formats like CSV, JSON, XML etc. preferably from S3 or other cloud storage.
  • Optimize SQL queries for performance and optimization in data warehouse technologies preferably Snowflake.
  • Implement DBT tests for data validation and quality checks.
  • Generate code using configurations with Python and Jinja templates.
  • Implement CI/CD process from scratch using GitHub.
  • Utilize AWS S3 for data storage, ECS, and IAM, including best practices for organization and security.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service