We are hiring an SRE to own reliability, performance, data freshness, and operational readiness for our production ETL platform. The platform supports data ingestion, transformation, and loading workflows across Kubernetes-based environments — including Airflow-based loader jobs and Spark-on-EKS jobs that load data into a Datalake/Lakehouse. You will operate and triage production pipelines end-to-end: extractors, loaders, batch jobs, streaming ingestion (Kafka), Spark workloads, and Airflow DAGs. You will tune Kubernetes and Spark for stability, build observability tooling, drive root cause analysis, write automation to reduce toil, and manage configuration and secrets through GitOps-style processes. The ideal candidate troubleshoots distributed systems from logs, metrics, and infrastructure signals, and drives permanent fixes through engineering partnership.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed