Sr Data Engineer

McGraw Hill LLC.UNAVAILABLE, UNAVAILABLE
Remote

About The Position

The Senior Data Engineer in Data and Analytics is responsible for advancing McGraw-Hill Education's (MHE) business intelligence and data platform capabilities, delivering scalable, reliable, and actionable insights across financial, product, customer, user, and third-party data domains. This role is deeply hands-on — designing, building, and optimizing end-to-end data pipelines and architectures on AWS (including services such as S3, Glue, Redshift, Lambda, EMR, and Step Functions) and Databricks (leveraging Delta Lake, Unity Catalog, and MLflow where applicable). The Senior Data Engineer will architect and implement dynamic reporting, analytics, and data modeling solutions that drive measurable outcomes in the education domain, while ensuring the performance, efficiency, and reliability of the broader Data Platform. The ideal candidate brings a strong data engineering foundation with deep, hands-on expertise in AWS cloud infrastructure and Databricks, including experience with Delta Lake architecture, medallion (Bronze/Silver/Gold) data design patterns, and Databricks Workflows for pipeline orchestration. Advanced proficiency in SQL and experience with Python or Scala for large-scale data transformation are essential. Familiarity with infrastructure-as-code (e.g., Terraform) and CI/CD practices for data pipelines is a strong plus. This role requires close collaboration with business stakeholders, data analysts, and product teams to translate complex data requirements into robust, production-grade engineering solutions — ensuring timely, high-quality delivery across all data initiatives. This is a remote position open to applicants authorized to work for any employer within the United States.

Requirements

  • Deep expertise in modern data lakehouse architecture, including Delta Lake, medallion design patterns, Unity Catalog governance, and the transition from traditional data warehousing to cloud-native lakehouse solutions on Databricks.
  • 5+ years of experience in Data Engineering.
  • Databricks — Delta Live Tables (DLT), Databricks Workflows, Unity Catalog, Delta Lake (MERGE, OPTIMIZE, VACUUM, Z-ordering), Databricks SQL, and MLflow.
  • AWS services — S3, Redshift, Glue, Lambda, EMR, Athena (with Iceberg), Step Functions, and IAM — integrated with Databricks as the primary compute and transformation layer.
  • Scripting and programming languages — Python (PySpark), Scala (Spark), or SQL as primary languages for pipeline development and data transformation within Databricks.
  • 3+ years of experience working with cloud platforms — primarily AWS — architecting and operating Databricks environments including workspace configuration, cluster policies, instance profiles, and cost optimization strategies.
  • 1+ years of experience with workflow automation and pipeline orchestration using Databricks Workflows, Apache Airflow (with the Databricks provider), or equivalent cloud-native schedulers, replacing traditional Unix shell scripting with scalable, observable pipeline management.
  • Advanced proficiency in SQL.
  • Experience with Python or Scala for large-scale data transformation.

Nice To Haves

  • Familiarity with infrastructure-as-code (e.g., Terraform) and CI/CD practices for data pipelines.
  • Experience with Publication and Education domain.
  • Prior experience or familiarity with Tableau/Alteryx.

Responsibilities

  • Design and deliver data solutions on Databricks, including building and maintaining lakehouses using Delta Lake with a medallion (Bronze/Silver/Gold) architecture.
  • Work with data from financial and operational systems, implementing Slowly Changing Dimensions (SCD Types 1, 2, and 3) using Delta Lake MERGE operations and Databricks SQL within a unified lakehouse model.
  • Run and optimize cloud data platforms on Databricks, including cluster configuration, autoscaling policies, job scheduling via Databricks Workflows, and adherence to daily runbook SLAs through proactive monitoring and alerting.
  • Utilize Git-based version control integrated into Databricks (Databricks Repos / Git folders) and project management tools such as Jira, operating within Agile/Kanban delivery frameworks.
  • Apply modern data architecture principles, including Unity Catalog for data governance, Delta Sharing, and cloud-native lakehouse design patterns on AWS with Databricks.
  • Translate business requirements into technical designs and deliver production-grade data solutions within Databricks, from initial scoping through deployment.
  • Design and develop parallel and distributed ETL/ELT pipelines using Apache Spark (PySpark/Scala) on Databricks, applying partitioning, caching, and broadcast join strategies for optimal resource efficiency and throughput.
  • Implement data mapping and transformation requirements using Databricks-native constructs including Spark transformations (aggregations, joins, unions, window functions, lookups, and pivot/unpivot operations) and Delta Live Tables (DLT) for declarative pipeline development.
  • Develop and maintain Databricks Workflows and job orchestration logic (including dependency management, retry policies, and alerting), replacing traditional shell-based wrapper patterns with cloud-native, maintainable pipeline automation.
  • Design and build integrations that support standard data modeling constructs — fact tables, dimension tables, star and snowflake schemas, and aggregations — implemented as Delta tables within Unity Catalog.
  • Provide end-to-end technical guidance across the full software development life cycle, from requirements gathering and architecture design through implementation, testing, and production deployment on Databricks.
  • Produce high-quality solution design documentation, including data flow diagrams, pipeline architecture specs, and Unity Catalog data asset definitions, ensuring clarity for both technical and business stakeholders.

Benefits

  • medical and/or other benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service