About The Position

We are looking for a part-time Freelancer for an AWS Data Engineer Position. This role requires relevant experience with a mandatory tech stack including Salesforce, AWS services, Terraform, and SQL. The primary focus will be on building and maintaining AWS data pipelines, optimizing data lake performance, managing modern data lake architecture, providing production support, and ensuring data quality and security.

Requirements

  • Source: Salesforce (ECRM & OSC)
  • Backup: Grax
  • Cloud: AWS (EC2, S3, CloudWatch, DynamoDB, RDS, Secrets Manager, ALB, ASG)
  • Infrastructure as Code: Terraform
  • File Format: Parquet
  • Target: Enterprise Data Lake (EDL)
  • Schema: Blue Schema
  • Database: Blue Database
  • Query Language: SQL
  • Develop ETL/ELT pipelines using AWS Glue, PySpark, Python, and SQL.
  • Ingest data from sources like Salesforce, databases, APIs, and S3.
  • Load curated data into Amazon Redshift and data lake storage.
  • Convert JSON/CSV data into Parquet.
  • Use Snappy compression.
  • Design proper partitioning strategies.
  • Resolve split limit and performance issues.
  • Optimize Athena query costs.
  • Work with Apache Iceberg tables.
  • Perform migrations from traditional Parquet tables.
  • Support schema evolution, time travel, and ACID transactions.
  • Investigate Glue jobs that suddenly become slow.
  • Debug Lambda timeouts.
  • Fix missing records and data quality issues.
  • Resolve Redshift performance problems.
  • Perform root cause analysis (RCA).
  • Build AWS infrastructure using Terraform.
  • Create reusable modules.
  • Manage Auto Scaling Groups, ALBs, IAM, Lambda, Secrets Manager, S3, Redshift, and DynamoDB.
  • Troubleshoot Terraform state and production deployment issues.
  • Manage AWS Secrets Manager.
  • Configure Lambda-based secret rotation.
  • Ensure Terraform does not overwrite rotated passwords.
  • Implement IAM least-privilege access.
  • Work with enterprise backup tools like Rubrik (their environment may use Grax for Salesforce).
  • Validate backup jobs.
  • Perform restores.
  • Support disaster recovery testing.
  • Verify restored data.
  • Validate source and target record counts.
  • Maintain audit/control tables (for example in DynamoDB).
  • Check checksums and duplicate records.
  • Troubleshoot data discrepancies reported by business users.
  • AWS Services: S3, Glue, Athena, Lambda, Redshift, DynamoDB, EventBridge, Step Functions, CloudWatch, Secrets Manager, IAM, Auto Scaling Groups, Application Load Balancer

Responsibilities

  • Build and maintain AWS data pipelines, developing ETL/ELT pipelines using AWS Glue, PySpark, Python, and SQL.
  • Ingest data from sources like Salesforce, databases, APIs, and S3.
  • Load curated data into Amazon Redshift and data lake storage.
  • Optimize Athena and data lake performance by converting JSON/CSV data into Parquet with Snappy compression, designing proper partitioning strategies, resolving split limit and performance issues, and optimizing Athena query costs.
  • Manage modern data lake architecture, working with Apache Iceberg tables, performing migrations from traditional Parquet tables, and supporting schema evolution, time travel, and ACID transactions.
  • Provide production support and troubleshooting, investigating slow Glue jobs, debugging Lambda timeouts, fixing missing records and data quality issues, resolving Redshift performance problems, and performing root cause analysis (RCA).
  • Build AWS infrastructure using Terraform, creating reusable modules, and managing Auto Scaling Groups, ALBs, IAM, Lambda, Secrets Manager, S3, Redshift, and DynamoDB.
  • Manage AWS Secrets Manager, configure Lambda-based secret rotation, ensure Terraform does not overwrite rotated passwords, and implement IAM least-privilege access.
  • Work with enterprise backup tools like Rubrik (or Grax for Salesforce), validate backup jobs, perform restores, support disaster recovery testing, and verify restored data.
  • Ensure data quality by validating source and target record counts, maintaining audit/control tables (e.g., in DynamoDB), checking checksums and duplicate records, and troubleshooting data discrepancies reported by business users.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service