The role focuses on owning the cost and performance of the inference stack, ensuring efficient model serving as workloads, traffic, and hardware evolve. The engineer will collaborate with fleet operators while managing core performance aspects like caching, batching, quantization, decoding, and kernel-level optimization. The goal is to enhance throughput and latency without sacrificing reliability or model quality.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed