Principal Software Engineer, Data Processing Optimization

NetApp•San Jose, CA
•$227,800 - $338,800

About The Position

DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing. DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.

Requirements

  • Bachelor’s degree in computer science. Master’s or PhD preferred, especially in areas related to database systems.
  • 15+ years of experience developing query planning and query optimization techniques for enterprise data processing software platform.
  • Prefer at least 5+ years of that experience on parallel platforms such as Apache Spark, Presto/Trino, or Apache Flink.
  • Deep knowledge of SQL semantics and expertise in SQL processing.
  • Additional Python processing preferred.
  • Solid understanding of databases, data warehouses, query engines.
  • Deep knowledge of query optimization state of the art, from published literature and open-source software.
  • Prefer own open-source contributions and/or publications in this area.
  • Solid understanding of the architecture, deployment, and operations of data processing platforms in applications such as data engineering, data preparation, and analytics.
  • Demonstrated proficiency in all phases of software development and post-production support of enterprise software or SaaS.
  • Adept in adopting AI effectively.
  • Experience owning the definition and delivery of major product capabilities through technical leadership and collaboration spanning multiple teams.
  • Strong collaboration and communication skills.

Responsibilities

  • Research, design, implement, and deliver logical and physical query plan optimization techniques for massively parallel data processing engines.
  • Develop query planning and query optimization techniques that enable industry-leading performance and performance efficiency on execution engines based on heterogeneous accelerated compute elements.
  • Integrate DataPelago data processing acceleration and plan optimization capabilities with diverse open-source engines including Apache Spark and Apache Flink.
  • Ensure product performance on customers’ real deployment environment and workload.
  • Establish metrics for plan quality, workload performance, reliability and lead the team in ensuring these metrics are continuously improved and achieved in customer deployments.
  • Provide technical leadership through all phases of product development and post-production support.
  • Own development and support of major components and capabilities.
  • Collaborate with other engineering team members and product management in defining product features, architecture, interfaces, and development timelines.
  • Mentor junior team members.

Benefits

  • Health Insurance
  • Life Insurance
  • Retirement or Pension Plans
  • Paid Time Off
  • various Leave options
  • Performance-Based Incentives
  • employee stock purchase plan
  • restricted stocks (RSU’s)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service