Data Engineer II

EchoStar•Centennial, CO
•Onsite

About The Position

This role addresses the complex challenges of constructing and optimizing high-performance data pipelines for both batch and streaming data processing across diverse internal and external sources. The position drives enterprise efficiency by leading critical Proof-of-Concept initiatives to evaluate and implement modern workflow orchestration tools and pioneering the integration of Generative AI services. Success means establishing reliable, scalable cloud infrastructure and robust automation that ensures data governance, quality, and optimal cost-effectiveness across the entire data lifecycle.

Requirements

  • Expert-level proficiency in Python, PySpark, and SQL alongside advanced Spark programming and monitoring of Big Data data engineering jobs
  • Deep practical experience leveraging AWS Big Data services, specifically Amazon EMR, EC2, S3, and modern workflow orchestrators like Apache Airflow or Control M
  • Proficiency with Git and GitLab for version control and CI/CD pipeline implementation alongside Infrastructure-as-Code tools like Terraform or AWS CloudFormation
  • AI Literacy and Innovation through the application of Generative AI tools like Amazon Q to optimize data engineering workflows, automated quality checks, and pipeline documentation
  • Strong analytical, problem-solving, and communication skills to effectively present findings from Proof-of-Concept initiatives and collaborate within Agile teams
  • Minimum Education: Bachelor’s Degree in Computer Science, Data Engineering, or a related technical field
  • Minimum Experience: 2 years of experience in Big Data and Data Engineering
  • Required Technical Skills: Must have at least 2 years of experience with: Python and SQL
  • Amazon EMR and Apache Spark
  • Apache Airflow or Control M
  • GitLab CI/CD pipelines
  • Amazon Q or Generative AI tools

Nice To Haves

  • Experience with Databricks preferred

Responsibilities

  • Design, construct, and optimize scalable data pipelines for batch and streaming data processing from various internal and external sources
  • Develop and manage data processing jobs using Apache Spark on Amazon EMR clusters while ensuring performance, cost-efficiency, and scalability
  • Implement transformation logic and complex data workflows primarily using Python, PySpark, and SQL
  • Lead a Proof-of-Concept project evaluating Apache Airflow as the enterprise-wide workflow orchestration tool, designing and deploying Directed Acyclic Graphs to manage dependencies
  • Explore and implement integration points for Amazon Q into data pipelines as part of the orchestration POC for automated data quality checks, data documentation generation, or pipeline optimization
  • Implement CI/CD pipelines for data platform components using GitLab, utilizing Infrastructure-as-Code templates and maintaining strict data governance, security, and quality throughout the lifecycle

Benefits

  • Versatile health perks
  • Flexible spending accounts
  • HSA
  • 401(k) Plan with company match
  • ESPP
  • Career opportunities
  • Flexible time away plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service