We are seeking an experienced Inference Engineer to build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics. You will design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization. This role involves implementing efficient low-level code (CUDA, Triton, custom kernels) and integrating it seamlessly into high-level frameworks. You will optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation), and develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed