Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work. In this role, you'll own token processing down to the lowest layers of the stack. You'll do things like: develop a new request scheduling strategy, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed