Site Rel Engineer Sr. - Model Integration Platform

PNCStrongsville, OH
$86,250 - $158,125Onsite

About The Position

As a site reliability engineer within PNC's Information Technology Group and Site Reliability Center (SRC), you will be based at one of PNC's Information Technology Hubs. This role is responsible for identifying and establishing ways of stabilizing environments and sites while assessing opportunities to drive engineering stability through analytics and metrics. You will be responsible for site design consulting, platform management, and capacity planning. Additionally, you will define, create, and ensure robust dashboards and monitoring systems are implemented, and identify, coordinate, and implement SLAs and SLOs through the use of established robust monitoring for applications, service sites, and platforms. You will lead in analyzing metrics from operating sites and applications to assist in performance tuning and fault finding, troubleshoot priority incidents, and participate in blameless post-mortems. You will understand the technology stack to optimize complex systems and user interactions, including code deployment, configuration, monitoring, availability, latency, change management, emergency response, and capacity planning of services in production. You will engage in testing strategy approaches and results, complex incident response and root cause analysis efforts, resolving underlying issues, driving continuous improvement in incident management processes, and reducing mean time to resolution. You will also mentor and train junior team members on best practices for infrastructure management and disaster recovery.

Requirements

  • Linux (RHEL) Administration
  • IBM GPFS / Spectrum Scale
  • Slurm Scheduler Administration
  • Shell/Bash Scripting
  • Basic Python
  • Linux Networking & Security
  • Storage/File Systems (LVM, NAS)
  • Spark / Jupyter
  • Production Support & Incident Resolution
  • Performance Monitoring & Troubleshooting
  • Bachelors degree or equivalent combination of education, job specific certification(s), and experience (including military service)
  • 3+ years of relevant / direct industry experience

Nice To Haves

  • IBM Spectrum Conductor
  • Cloud (AWS/Azure)
  • Ansible/Python Automation
  • VMware/KVM
  • Docker/Kubernetes
  • Databases (Oracle/MySQL/Postgres)
  • Tableau/Power BI Reporting
  • Apache Tomcat installation, tuning, and administration

Responsibilities

  • Identifies and establishes ways of stabilizing environments and sites while assessing opportunities to drive engineering stability through the analytics and metrics.
  • Responsible for site design consulting, platform management, and capacity planning.
  • Defines, creates, and ensures that robust dashboards and monitoring systems are implemented.
  • Identifies, coordinates, and implements SLAs and SLOs through the use of established robust monitoring for applications, service sites and platforms.
  • Leads in analyzing metrics from operating sites and applications to assist in performance tuning and fault finding.
  • Troubleshoots priority incidents and participates in blameless post-mortems.
  • Understands our stack to optimize complex systems and user interactions, such as how code is deployed, configured, and monitored, as well as the availability, latency, change management, emergency response, and capacity planning of services in production.
  • Engages in testing strategy approaches and results, complex incident response and root cause analysis efforts, resolving underlying issues, driving continuous improvement in incident management process, and reducing mean time to resolution.
  • Mentors and trains junior team members on best practices for infrastructure management and disaster recovery.

Benefits

  • medical/prescription drug coverage (with a Health Savings Account feature)
  • dental and vision options
  • employee and spouse/child life insurance
  • short and long-term disability protection
  • 401(k) with PNC match
  • pension and stock purchase plans
  • dependent care reimbursement account
  • back-up child/elder care
  • adoption, surrogacy, and doula reimbursement
  • educational assistance, including select programs fully paid
  • a robust wellness program with financial incentives
  • maternity and/or parental leave
  • up to 11 paid holidays each year
  • 9 occasional absence days each year, unless otherwise required by law
  • between 15 to 25 vacation days each year, depending on career level; and years of service
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service