This is an application to join the talent network of Dragonfly, a crypto-native Venture Capital and research firm. They are sourcing for a Senior Inference Optimization Engineer for one of their portfolio companies. This company is building privacy-first consumer AI infrastructure and is looking for someone to work on the bleeding edge of LLM inference performance, focusing on throughput, latency, and cost per token at scale. The role involves optimizing GPU infrastructure, benchmarking inference engines, optimizing load-balancing algorithms, evaluating new inference optimization techniques, and assessing emerging inference hardware.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed