As a site reliability engineer within PNC's Information Technology Group and Site Reliability Center (SRC), you will be based at one of PNC's Information Technology Hubs. This role is responsible for identifying and establishing ways of stabilizing environments and sites while assessing opportunities to drive engineering stability through analytics and metrics. You will be responsible for site design consulting, platform management, and capacity planning. Additionally, you will define, create, and ensure robust dashboards and monitoring systems are implemented, and identify, coordinate, and implement SLAs and SLOs through the use of established robust monitoring for applications, service sites, and platforms. You will lead in analyzing metrics from operating sites and applications to assist in performance tuning and fault finding, troubleshoot priority incidents, and participate in blameless post-mortems. You will understand the technology stack to optimize complex systems and user interactions, including code deployment, configuration, monitoring, availability, latency, change management, emergency response, and capacity planning of services in production. You will engage in testing strategy approaches and results, complex incident response and root cause analysis efforts, resolving underlying issues, driving continuous improvement in incident management processes, and reducing mean time to resolution. You will also mentor and train junior team members on best practices for infrastructure management and disaster recovery.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior