Parasail is redefining AI infrastructure by enabling seamless deployment across a distributed network of GPUs, optimizing for cost, performance, and flexibility. Our mission is to empower AI developers with a fast, cost-efficient, and scalable cloud experience—free from vendor lock-in and designed for the next generation of AI workloads. The Senior/Staff Inference Reliability Engineer will own the end-to-end reliability and production performance of customer inference workloads. This role sits at the intersection of inference platform engineering, LLM performance, and infrastructure reliability. You will ensure that customer endpoints meet expectations for availability, latency, throughput, quality, and cost. When an endpoint degrades, you will follow the problem across the entire serving path—from APIs, routing, scheduling, and autoscaling through model servers, GPUs, networking, and underlying infrastructure—and drive it through resolution. This is not a traditional DevOps role focused only on clusters and deployments. It is a production systems role for an engineer who enjoys investigating ambiguous performance problems, building diagnostic tooling, and turning recurring incidents into durable platform improvements. Prior LLM-inference experience is valuable but not required. We are looking for someone with deep production systems experience who can quickly learn inference-specific technologies and metrics.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed