Senior Data Reliability Engineer AWS

Empower•Overland Park, KS
•$105,700 - $149,275•Hybrid

About The Position

We are seeking a hands-on Senior Data Reliability Engineer to ensure the reliability, stability, and performance of our AWS-based data platform. You will troubleshoot production data systems, resolve incidents, and partner with engineering teams to meet business-critical SLAs and continuously improve operational practices. This role requires a strong Data Engineering foundation combined with production reliability experience. The ideal candidate has personally built data pipelines, supported them in production, diagnosed real-world data and processing failures, and implemented improvements to prevent recurrence.

Requirements

  • 5+ years of hands-on experience building, operating, or supporting production data platforms, with significant recent experience in AWS environments.
  • Hands-on experience designing, building, deploying, and operating production-grade data pipelines.
  • Prior experience building data pipelines and seeing them through production, including exposure to real-world failures and operational challenges.
  • Strong hands-on Python and SQL experience in production data environments.
  • Hands-on experience troubleshooting distributed data-processing systems such as Spark/EMR.
  • Experience working with AWS data services such as EMR, S3, Glue, Redshift, DynamoDB, Lambda, or similar data services.
  • Experience operating or troubleshooting cloud data warehouses such as Redshift, Snowflake, or similar platforms.
  • Strong understanding of end-to-end production data architecture, including ingestion, processing, orchestration, storage/warehousing, data quality, monitoring, and downstream consumption.
  • Experience handling production incidents and performing root cause analysis.
  • Experience with data validation, reconciliation, backfills, reprocessing, and late or incomplete data.
  • Strong problem-solving skills and the ability to work through ambiguous production data issues.
  • Strong communication during incidents with both technical and non-technical stakeholders.

Nice To Haves

  • Experience troubleshooting complex Spark/EMR production workloads.
  • Experience improving observability and alerting specifically for data pipelines and data platforms.
  • Experience with streaming or event-driven data systems such as Kafka, Kinesis, or CDC patterns.
  • Experience with Snowflake performance, workload, reliability, or cost optimization.
  • Experience designing automated controls for data freshness, completeness, reconciliation, and anomaly detection.
  • Experience with disaster recovery, backup validation, and resiliency testing for data platforms.
  • Experience developing automation or intelligent tooling to improve data-platform reliability, troubleshooting, performance, or cost efficiency.

Responsibilities

  • Own production data pipeline and platform reliability, resolving failures, delays, data quality issues, and performance problems.
  • Troubleshoot distributed data-processing workloads across Spark/EMR, ingestion pipelines, and data warehouses such as Redshift and Snowflake.
  • Lead or support incident response, perform root cause analysis, restore data safely, and implement lasting fixes.
  • Define and monitor data SLAs for freshness, latency, and completeness and improve monitoring, alerting, and observability.
  • Diagnose data issues including late-arriving data, incomplete loads, duplicate data, schema changes, and reconciliation failures.
  • Perform safe backfills, reprocessing, and production data recovery while preventing duplicate or inconsistent data.
  • Develop automation and tooling using Python and SQL to reduce operational toil, improve troubleshooting, and increase platform resilience.
  • Identify and address data-platform performance and cost issues.
  • Partner with Data Engineering teams to improve pipeline design, data quality, reliability, and production readiness.
  • Support disaster recovery, backup validation, and recovery workflows.
  • Create and maintain runbooks, SOPs, and operational documentation.
  • Participate in an on-call rotation for production data systems.

Benefits

  • Medical, dental, vision and life insurance
  • Retirement savings – 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup
  • Tuition reimbursement up to $5,250/year
  • Business-casual environment that includes the option to wear jeans
  • Generous paid time off upon hire – including a paid time off program plus ten paid company holidays and three floating holidays each calendar year
  • Paid volunteer time — 16 hours per calendar year
  • Leave of absence programs – including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA)
  • Business Resource Groups (BRGs) – BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are open to all.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service