This role focuses on optimizing large language models (LLMs) for production environments, making them faster, cheaper, and more reliable. The engineer will own the end-to-end inference stack, employing modern optimization techniques and delving into serving code when necessary. This is a hands-on engineering position involving core systems and performance work on demanding models. The role is applied, with optimizations directly impacting customer deployments. It requires tailoring deployments to customer needs, managing workloads from proof of concept to production, and ensuring performance gains are realized. The position involves coding, profiling, low-level optimization, and a customer-facing aspect, including product and technical solutions work.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior