AI Infrastructure Engineer

IntelHillsboro, CA
Hybrid

About The Position

We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance. You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang. Your optimizations will be instrumental in unlocking the full potential of Intel hardware for state-of-the-art generative AI workloads.

Requirements

  • Bachelors Degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years experience, Masters Degree and 3+ years, OR PhD.
  • 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC).
  • Proficiency in modern C++ and Python. You are comfortable reading and modifying complex systems-level code.

Nice To Haves

  • Understanding of CPU/GPU architecture.
  • Understanding of modern LLM architectures and inference paradigms: attention mechanisms, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation.
  • Prior open-source contributions to inference engines (vLLM, SGLang, PyTorch, llama.cpp).
  • Hands-on experience writing and optimizing custom GPU kernels using Triton, SYCL, CUDA/CUTLASS, or other DSLs.
  • Experience with scale-out inference orchestration across multi-node topologies.
  • You leverage AI coding agents daily to accelerate your own workflow and benchmark generation.

Responsibilities

  • Drive Inference Performance: Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs.
  • Deep Stack Optimization: Profile, diagnose, and resolve cross-stack performance bottlenecks.
  • Kernel Development and Integration: Design, write, and optimize custom high-performance kernels for critical attention mechanisms, MoE, quantization, and operator fusions.
  • Open Source Leadership: Upstream your architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch, acting as a bridge between the hardware teams and the open-source community.
  • Shape the Hardware Roadmap: Apply roofline analysis and systematic profiling to decompose bottlenecks. You will partner with our architecture and compiler teams to shape future GPU roadmaps based on real-world GenAI workload data.
  • Show passion about AI infrastructure and performance optimization.

Benefits

  • competitive pay
  • stock bonuses
  • health
  • retirement
  • vacation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service