We are seeking an experienced AI Inference Engineer to help design, deploy, and operate private AI inference infrastructure at rack scale. This role is focused on the engineering of a production-grade inference platform: architecture, serving stack design, benchmarking, integration, reliability, and operational readiness. The ideal candidate has deep hands-on experience building AI inference systems and the platform layers that support them, including multi-GPU inference, Kubernetes-based deployment, autoscaling, observability, routing, access control, and production support.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager