Designs, develops, and optimizes production Data Products on the Databricks Lakehouse Platform running on AWS, with reliability, observability, performance, scalability, and cost efficiency engineered directly into the solutions built. Engineers monitoring, logging, metrics, dashboards, and proactive alerting to improve diagnostics and production visibility. Identifies recurring reliability issues, performs root cause analysis for complex production issues, and automates recovery and operational activities where practical. Optimizes Spark workloads, Delta Lake, Databricks compute, SQL Warehouses, and storage, diagnosing performance bottlenecks and engineering sustainable improvements while continuously optimizing Databricks and AWS cost efficiency. Develops reusable observability and production-readiness patterns and recovery procedures, supporting appropriate resilience and business continuity requirements. Codes production-grade ETL/ELT and data-processing pipelines using Python/PySpark, SQL, Spark, and Delta Lake, following approved data models and Medallion Architecture patterns. Leads technical reviews and mentors engineers on Databricks development, observability, troubleshooting, and performance optimization. Note: May be internal or external-facing, working in conjunction with Platform Engineering, Governance Engineering, and enterprise architecture functions. May include enterprise-wide, AI-enabled data solutions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior