Senior Software Engineer – Foundational Data Systems for AI

GranicaSan Francisco, CA
$190,000 - $250,000Onsite

About The Position

AI today is limited not only by model design but by the inefficiency of the data that feeds it. At scale, each redundant byte, each poorly organized dataset, and each inefficient data path slows progress and compounds into enormous cost, latency, and energy waste. Granica’s mission is to remove that inefficiency. We combine new research in information theory, probabilistic modeling, and distributed systems to design self-optimizing data infrastructure: systems that continuously improve how information is represented and used by AI. This engineering team partners closely with the Granica Research group led by Prof. Andrea Montanari (Stanford), bridging advances in information theory and learning efficiency with large-scale distributed systems. Together, we share a conviction that the next leap in AI will come from breakthroughs in efficient systems, not just larger models.

Requirements

  • Depth in distributed systems: consensus, partitioning, replication, fault tolerance.
  • Experience with columnar formats such as Parquet or ORC and low-level encoding strategies.
  • Understanding of metadata-driven architectures and adaptive query planning.
  • Production experience with Spark, Flink, or custom distributed engines on cloud object storage.
  • Proficiency in Java, Rust, Go, or C++ with an emphasis on clarity and quality.
  • Curiosity about theory of the mathematics of compression, entropy, and learning efficiency.
  • A builder’s mindset: pragmatic, rigorous, and grounded in long-term systems thinking.

Nice To Haves

  • Familiarity with Iceberg, Delta Lake, or Hudi.
  • Research or open-source contributions in compression, indexing, or distributed computation.
  • Interest in how data representation affects training dynamics and model reasoning efficiency.

Responsibilities

  • Architect the transactional and metadata substrate that supports time-travel, schema evolution, and atomic consistency across petabyte-scale tabular datasets.
  • Build systems that reorganize data autonomously, learning from access patterns and workloads to maintain peak efficiency without manual tuning.
  • Optimize bit-level organization (encoding, compression, layout) to extract maximal signal per byte read.
  • Develop distributed compute systems that scale predictively, adapt to dynamic load, and maintain reliability under failure.
  • Implement new algorithms in compression, representation, and optimization emerging from ongoing research. Opportunities to publish and open-source are encouraged.
  • Design for minimal time between question and insight, enabling models and humans to learn faster from data.

Benefits

  • Competitive salary, meaningful equity, and performance bonus for top performers
  • 401(k) with company match, comprehensive health coverage, and unlimited PTO
  • Daily catered meals in our Mountain View office
  • Support for research, publication, and conference participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service