At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research. We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models. About the Role We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture. You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations). You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior