Lavendo partners with startups and high‑growth companies to help them hire top‑tier sales, GTM, and technical talent. This role is with one of our clients; we’ll share full details about the company and interview process as we get to know you and confirm mutual fit. Our client is building the autonomous performance engineer the AI industry doesn't have enough of: AI agents that optimize GPU kernels so companies can run inference on open-source LLMs faster and cheaper than anyone else. Their product wins on the two numbers that matter most to engineering buyers — tokens/second and latency — and on cost, backed by GPU procurement relationships most seed-stage companies don't have access to. They're a YC company from a recent batch, having just closed a $4M seed round led by a respected early-stage fund. The angel list is the kind you don't usually see this early: senior technical leaders from Google, OpenAI, Dropbox, and Together AI have personally put money in. If you've spent any time around inference or GPU infra, that's not a vanity signal — it's some of the most technically demanding people in the space betting on this team's approach. Under 10 people today, real customers running production inference, and a founder who still closes every deal himself. This is about as early as 'early-stage' gets. The Mission Make it possible for any team — not just the ones with in-house GPU experts — to serve state-of-the-art open-source models in production at a fraction of the usual latency and cost. Less time hand-tuning kernels, more time shipping product.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed