The purpose of this role is to resolve priority incident tickets and service requests while contributing to incident analysis, change request preparation, and continuous service improvement. This role involves performing initial triage for high priority incident tickets (P2 tickets) which may have an impact across one or more processes and resolving priority incident (P3) tickets within defined SLAs. It also includes fulfilling software/hardware/network service requests within agreed timelines, monitoring SLA timelines for the entire lifecycle of high priority incident tickets, creating change requests, and analyzing incidents to support root cause analysis and related service improvement plans. The role requires a systematic problem-solving approach, strong communication skills, a sense of ownership, and the drive to find solutions. Additionally, it involves working closely with the development team on maintaining the operational health of core compute services for API availability and low latency, managing and triaging tickets, driving prioritization and execution of work based on impact, and scaling systems sustainably through mechanisms such as easy-to-use tooling and automation. The role also focuses on evolving systems/products for better scalability, reliability, and development velocity, driving new runbooks to help reduce mean triage time of incidents, and prioritizing and automating high hit count runbooks, while practicing sustainable incident response and driving root cause analysis.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed