CoreWeave’s Storage Reliability team sits at the intersection of infrastructure engineering, operations, and customer enablement. The team is responsible for ensuring the stability, performance, and operational excellence of the storage systems powering some of the world’s largest AI workloads. We work directly with production systems at scale, partnering closely with engineering, solutions, and customer-facing teams to maintain reliability while continuously improving the tooling, automation, and observability that support our storage platform. About the role: As a Storage Reliability Engineer, you will operate and support mission-critical storage systems that power large-scale AI and data-intensive workloads. You will work hands-on with production infrastructure, triaging complex incidents, debugging issues across the application, system, and kernel layers, and contributing fixes and improvements to the storage stack. This role sits at the boundary between engineering and operations, turning real-world production learnings into long-term reliability improvements through tooling, automation, and operational best practices. You’ll also partner closely with internal teams and customers to diagnose and resolve complex deployment and performance issues.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level