We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance. You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang. Your optimizations will be instrumental in unlocking the full potential of Intel hardware for state-of-the-art generative AI workloads.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Ph.D. or professional degree