Runpod is the AI Developer Cloud, serving over a million developers who use the platform to experiment, train, fine-tune, deploy, and scale AI. Having processed more than 20 billion inference requests and recently closing a $100M Series A, Runpod is at a pivotal moment in AI infrastructure. The company is seeking individuals who are passionate, driven, and aim to make a significant impact at scale. This is a remote-first role within a small, agile team that values ownership and speed. This role focuses on making Runpod the premier platform for LLM inference, prioritizing speed and cost-efficiency. The engineer will lead efforts to optimize LLM serving performance end-to-end, encompassing measurement, analysis, and improvement across various models, hardware, and workloads. The work will directly influence customer experience regarding latency and cost. It's a hands-on engineering position for someone adept at identifying and resolving performance bottlenecks, and implementing robust production solutions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed