The purpose of this role is to assist in the triage and documentation of priority incidents, support resolution of software and hardware service requests and ensure SLA adherence by escalating delays when required. The role involves working closely with the development team on maintaining the operational health of core compute services for API availability and low latency, managing and triaging tickets, driving prioritization and execution of work based on impact, and scaling systems sustainably through mechanisms such as easy-to-use tooling and automation. The individual will work in concert with service developers to evolve systems/products for better scalability, reliability, and development velocity, drive new runbooks to help reduce mean triage time of incidents, prioritize and automate high hit count runbooks, and practice sustainable incident response and drive root cause analysis.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed