About The Position

Sonatype is the software supply chain security company. We provide the world’s best end-to-end software supply chain security solution, combining the only proactive protection against malicious open source, the only enterprise grade SBOM management and the leading open source dependency management platform. This empowers enterprises to create and maintain secure, quality, and innovative software at scale. As founders of Nexus Repository and stewards of Maven Central, the world’s largest repository of Java open-source software, we are software pioneers and our open source expertise is unmatched. We empower innovation with an unparalleled commitment to build faster, safer software and harness AI and data intelligence to mitigate risk, maximize efficiencies, and drive powerful software development. More than 2,000 organizations, including 70% of the Fortune 100 and 15 million software developers, rely on Sonatype to optimize their software supply chains. We’re looking for a Staff Data Engineer to join our growing Data Platform team. You’ll play a key role in designing and scaling the infrastructure and pipelines that power analytics, machine learning, and business intelligence across Sonatype.You’ll work closely with stakeholders across product, engineering, and business teams to ensure data is reliable, accessible, and actionable. This role is ideal for someone who thrives on solving complex data challenges at scale and enjoys building high-quality, maintainable systems. At Sonatype, we: Use data with purpose: you'll get the chance to work on problems that directly impact how the world builds secure software Use modern tooling: you'll get the chance to leverage the best of open-source and cloud-native technologies Have a deep collaborative culture: you'll be joining a passionate team that values learning, autonomy, and impact

Requirements

  • 8+ years of experience as a Data Engineer or in a similar backend engineering role
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field
  • Databricks Optimization: Tune Spark jobs, optimize join performance, and manage Delta Lake architecture for batch and streaming data.
  • Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality and ops.
  • Strong programming skills in Python, Scala, or Java
  • Hands-on experience with distributed data systems like Spark or Kafka
  • Proficient in writing complex SQL and NoSQL queries and optimizing queries for performance
  • Experience building and maintaining robust ETL/ELT pipelines in production
  • Understanding of data modeling techniques (star schema, dimensional modeling, etc.)

Nice To Haves

  • Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data
  • A track record of improving data platform reliability, scalability, performance, and cost efficiency
  • Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)
  • Hands-on experience with cloud data platforms, particularly AWS
  • Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi
  • Experience implementing data observability, lineage, governance, and automated data quality frameworks
  • Experience designing real-time or streaming data architectures using data lake technologies

Responsibilities

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes
  • Architect and optimize data models and storage solutions for analytics and operational use
  • Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets
  • Own and evolve parts of our data platform using Databricks and Spark
  • Implement observability, alerting, and data quality monitoring for critical pipelines
  • Drive best practices in data engineering, including documentation, testing, and CI/CD
  • As a Staff Engineer you will help drive long-term architectural vision and mentor the team on engineering best practices, while partnering with stakeholders to ensure data solutions support business outcomes.
  • Contribute to the design and evolution of our next-generation data lakehouse architecture

Benefits

  • Parental Leave Policy
  • Paid Volunteer Time Off (VTO)
  • flexible working practices
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service