Sparibis LLC is seeking a Data Engineer to provide expertise in the development, implementation, integration, and sustainment of data architectures, data hubs, data lakes, and data warehouse solutions. The role involves supporting the design and development of data models, data structures, and data acquisition processes to enable data-driven decision making across the HC/HR Data Domain. The Data Engineer will develop engineering and implementation plans for data hubs, data acquisition, data modeling, data integration, and related data engineering activities. This includes researching and evaluating existing data sources to identify authoritative sources for data hub development, planning and maintaining data architectures, and designing, developing, maintaining, and optimizing ETL/ELT data pipelines and data transformation processes. The position requires developing and maintaining batch and streaming data pipelines using technologies such as Spark, Python, Databricks, Palantir Foundry, Kafka, and related data engineering tools. Additionally, the role involves developing and maintaining data acquisition processes, implementing incremental data loading strategies, defining and implementing approaches for handling late-arriving data, and identifying opportunities to automate manual data processes. The Data Engineer will develop, maintain, and optimize data engineering solutions within Databricks and enterprise data lake environments, including configuring, monitoring, and managing Databricks clusters. Developing and maintaining Spark-based data processing solutions using PySpark, Spark SQL, Spark Data Frame, Data Sets, and related technologies, as well as implementing and maintaining Delta Lake solutions, is also a key responsibility. The role includes supporting data engineering and application development within Palantir Foundry, developing and maintaining streaming data solutions using Kafka, Kafka Streams, ksqlDB, and related technologies, and developing and maintaining Python-based data processing applications and AWS Lambda functions. Support for data integration and processing using AWS services such as S3, Kinesis, Lambda, and DynamoDB, along with monitoring, troubleshooting, and optimizing cloud-based data processing solutions, is required. The position also involves developing and maintaining data quality controls, validation processes, and data quality gates, developing and executing data-driven testing and unit testing, and establishing and maintaining data lifecycle policies and processes. Monitoring and troubleshooting data pipelines and processing jobs, and identifying opportunities to improve data processing performance, reliability, scalability, and maintainability are also key aspects of the role.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior