The purpose of this role is to assist in the triage and documentation of priority incidents, support resolution of software and hardware service requests and ensure SLA adherence by escalating delays when required. This role involves working closely with the development team on maintaining the operational health of core compute services for API availability and low latency, managing and triaging tickets, driving prioritization and execution of work based on impact, and scaling systems sustainably through mechanisms such as easy-to-use tooling and automation. The role also involves working in concert with service developers to evolve systems/products for better scalability, reliability, and development velocity, driving new runbooks to help reduce mean triage time of incidents, prioritizing and automating high hit count runbooks, and practicing sustainable incident response and driving root cause analysis.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed