The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. This position focuses on enhancing system reliability, scalability, and performance across high-density hardware and software infrastructure in regional data centers. Responsibilities include defining KPIs, proactive monitoring, automation, and urgent issue resolution. Collaboration with teams ensures best practices, reduced downtime, and optimized systems. The role supports seamless operations, business-critical applications, and improved user experiences through efficient, data-driven solutions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior