We are seeking a highly experienced Lead PySpark Engineer with over 10 years of experience in big data and distributed computing. The ideal candidate will have very strong hands-on experience with PySpark, Apache Spark, and Python. A strong command of SQL and NoSQL databases (DB2, PostgreSQL, Snowflake, etc.), proficiency in data modeling and ETL workflows, and familiarity with workflow schedulers like Airflow are essential. Experience with AWS cloud-based data platforms is required. This role involves leading the design, development, and deployment of PySpark-based big data solutions, architecting and optimizing ETL pipelines, and collaborating with various teams to deliver scalable solutions. The position also requires optimizing Spark performance, implementing data engineering best practices, and ensuring data security and compliance. Mentoring junior developers and code review are also key responsibilities.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed