Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence

NVIDIASanta Clara, WA
$184,000 - $356,500

About The Position

As part of NVIDIA’s Analytics and Data Intelligence (ADI) group, this team develops “libcudf”—the open-source CUDA C++ library that accelerates database and DataFrame operations. With flexible I/O, blazing-fast merging, aggregating, and filtering, our library serves diverse domains, including business intelligence, genomics, LLM training, and more. We use the latest tools in modern C++ and CUDA to produce software with elegant design, broad feature coverage, and best-in-class performance. The team is looking for an outstanding engineer and/or scientist to apply their parallel programming skills to accelerate open-source software libraries for GPU-based data processing. In this position, you will drive speed-of-light performance in structured data processing, spanning hardware from single workstations to multi-node GPU supercomputers. In addition, you will be building the computational core for DataFrame and database accelerators—highly optimized C++ and CUDA libraries that leverage the parallel nature of GPUs to accelerate operations from data loading and parsing, joins, aggregations, and more. Come bring your inspiration and problem-solving skills to our open-source software suite, and you can be our next major contributor!

Requirements

  • 8+ years of experience in Computer Science or Software Engineering
  • MS degree or PhD in computer science, engineering, or a related field, or equivalent experience
  • Strong Modern C++ programming skills
  • You care deeply about robust, readable, high-performance code

Nice To Haves

  • Expertise in high-performance communication protocols (e.g., UCX, NCCL) and distributed algorithms for query engines
  • Familiarity with RAPIDS cuDF
  • Experience in distributed workflow development and debugging
  • Passion for open-source software development and publishing your work in technical blogs and conferences

Responsibilities

  • Own development for “UcxExchange” in Velox (30%)
  • GPU-to-GPU communication in Presto: https://github.com/facebookincubator/velox/tree/main/velox/experimental/ucx-exchange
  • Drive the feature roadmap, improve performance, and ensure correctness
  • Optimize multi-node performance for analytical queries with Presto GPU (30%)
  • Targeting both on-prem clusters and public cloud servers
  • Lead projects of interest in Presto, Velox, cuDF (40%)
  • Such as operator completeness, GPU memory oversubscription, efficient CPU fallback
  • Plus your ideas

Benefits

  • competitive salaries
  • generous benefits package
  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service