Member of Technical Staff Inference performance depends on how efficiently models use the underlying hardware. Techniques across kernels, compilers, runtimes, and serving systems can dramatically improve latency and throughput, but these optimizations are difficult, hardware-specific, and slow to reproduce across new accelerators. And none of it counts until it is running in a customer's production traffic. At Wafer, we are building AI systems that automatically optimize inference workloads across silicon. The goal is fungible token capacity. Any accelerator optimized toward serving inference most efficiently. Wafer is well funded and serves trillions of tokens a month for mission critical workloads. We serve the highest performance inference to fast-growing AI startups. Members of Technical Staff build the systems that make that possible and own the customers running on them. There is no separate solutions team, and no layer between you and the workload.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed