About The Position

As a Principal Software Engineer on the Data Platform, you will own the architecture of the lakehouse engine that powers Innovaccer's platform in on-premise deployments: Apache Iceberg as the table format, Trino as the query engine, a REST catalog service, and Spark as transform compute. Cloud data warehouses have no on-premise equivalent, so this is a ground-up engine design, not a re-point. It is the single longest-lead technical track in the program, and the decisions you make on catalog, engine placement, and pipeline redesign gate everything downstream: transforms, serving, and reporting.

Requirements

  • B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.
  • 12+ years of industry experience building and operating large-scale data platforms or distributed systems.
  • Deep, hands-on expertise with distributed SQL engines: Trino/Presto or Spark SQL internals, query planning, and performance engineering.
  • Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg): table spec, merge-on-read versus copy-on-write, and table maintenance at scale.
  • Working knowledge of Iceberg catalog services (REST catalogs such as Polaris or Nessie, or Hive Metastore) and S3 compatible object storage.
  • Strong understanding of cloud warehouse internals (Snowflake, BigQuery, or Redshift) sufficient to design functional equivalents on open-source infrastructure.
  • Professional software development experience with Java and/or Python.
  • Experience delivering data platforms in on-premise, regulated, or air-gapped environments is a strong plus; healthcare data experience is a plus.

Responsibilities

  • Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout.
  • Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores.
  • Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators).
  • Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup.
  • Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL.
  • Mentor senior engineers across data workstreams, review designs, and raise the bar on engineering quality.
  • Partner with platform engineering on storage sizing, resource isolation, and capacity planning for the lakehouse footprint.

Benefits

  • Generous Paid Time Off: Recharge and relax with 20 days of fixed time off per year, in addition to company holidays—because we believe work-life balance fuels performance.
  • Best-in-Class Parental Leave: Spend quality time with your growing family. We offer one of the industry’s most generous parental leave policies to support you during life’s most important moments.
  • Recognition & Rewards: We celebrate wins—big and small. Get rewarded with monetary incentives and company-wide recognition for your impact and dedication. Your hard work won’t go unnoticed.
  • Comprehensive Insurance Coverage: Stay covered with medical, dental, and vision insurance, plus 100% company-paid short- and long-term disability and basic life insurance. Optional perks include discounted legal aid and pet insurance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service