Staff Engineer| Hybrid: Chicago, IL

Circana,
$170,000Hybrid

About The Position

We are seeking Staff Engineers to drive design, evolution, and execution of our next-generation, massively parallel processing (MPP) clustered database engine (Whitebox) that currently powers large-scale retail and consumer analytics solutions across Circana clients in 22+ countries and serving millions of queries every day in real-time against massive data sets. This is a highly technical role for an expert who bridges the gap between High-Performance Computing (HPC) and Big Data systems engineering and low-latency machine learning infrastructure and involves database internals, distributed systems, performance engineering, data processing at scale. In this role, you will be one of the primary designers and developer responsible for building low-latency, high-throughput distributed database kernels. You will design bare-metal optimized software layers, ensuring our database product fully exploits modern multi-core, vectorized hardware and distributed network fabrics to process terabyte/petabyte-scale datasets and where appropriate, seamlessly executes highly performant analytics and in-database ML inference directly on relational and the occasional vector datasets.

Requirements

  • Minimum - 4 years Computer Science engineering degree from a reputed institute. Masters in HPC (high performance computing) and Bigdata engineering areas preferred for Senior staff engineer and above roles
  • 6+ years of production experience writing hands-on, low-level systems in C++ (C++17/20/23 preferred). Mastery of template metaprogramming, concepts, coroutines, and custom memory allocators is key.
  • Very good to expert-level knowledge of cluster-scale distributed programming using MPI (Message Passing Interface) paired with shared-memory thread-level programming via OpenMP and MapReduce frameworks.
  • Documented experience with microarchitectural optimization, loop vectorization, pointer-aliasing constraints, and manual SIMD intrinsics programming.
  • Deep understanding of OS internals, memory barriers, lock-free data structures, and atomics (std::atomic).
  • Practical experience designing core database subsystems (e.g., vectorized query operators, LSM-trees, B-trees, custom buffer pools, columnar serialization frameworks and query plan optimizers).
  • Strong debugging, performance tuning, and problem-solving skills.
  • Proficiency in one or more leading memory and compute monitoring/performance tools (such as Valgrind, Purify, Intel/AMD/ARM diagnostics etc.)

Nice To Haves

  • Track record contributing to open-source or work experience on proprietary distributed engines (e.g., ClickHouse, DuckDB, Doris, CockroachDB, RocksDB, ScyllaDB, or Velox).
  • Experience with cloud-native distributed, elastic environments, object storage tiering (S3/GCS), or zero-copy networking.
  • Experience integrating heterogeneous acceleration (NVIDIA CUDA / ROCm) into big data workloads and hybrid CPU/GPU setups.
  • Experience building or modifying high-performance vector index/search engines, graph-based indices, or localized quantization frameworks.
  • Functional domain awareness – CPG, Retail, Pharma verticals

Responsibilities

  • Play a key role in core architectural choices and roadmap for our distributed database kernel, including query execution engines, custom/hybrid columnar storage layers, distributed transaction managers, and cluster coordination protocols.
  • Write, optimize, and maintain highly performant, production-ready backend code utilizing modern C++ (C++20/23).
  • Architect MPP execution engines that distribute, schedule, and execute queries across several dozens to hundreds of nodes using MPI and hybrid OpenMP models.
  • Profile and eliminate execution bottlenecks by implementing explicit SIMD vectorization (AVX-512, ARM Neon, SVE) and ensuring absolute cache locality (L1/L2/L3).
  • Design ultra-low latency cluster communication subsystems utilizing RDMA, RoCE, or kernel-bypass tech (DPDK).
  • Architect and implement a native, highly performant, zero-copy ML execution framework directly inside the MPP engine to eliminate data serialization and transport bottlenecks between the database and external AI pipelines.
  • Develop mathematically rigorous, cache-oblivious, and highly vectorized algorithms for real-time statistical computations, matrix operations/manipulations on clustered data
  • Design and optimize or embed/reference distributed high-dimensional vector storage subsystems, custom SIMD-accelerated indexing structures (e.g., HNSW, IVF-PQ), and parallelized approximate nearest neighbor (ANN) search algorithms.
  • Be comfortable working in a cross country/multi time zone team of elite systems engineers, implementing as well as conducting rigorous peer design and development reviews
  • Collaborate with cross-functional teams in defining and developing new platform capabilities.
  • Actively participate in architecture reviews, code reviews, and technical mentoring.
  • Investigate and resolve complex performance and scalability issues.

Benefits

  • Highly competitive compensation package (Base, Bonus)
  • Stimulating work environment and a great Circana work culture that allows members to bring their best daily.
  • paid time off
  • medical/dental/vision insurance
  • 401(k)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service