Sr. Specialist, SRE – Compute Platforms

Canadian Tire CorporationToronto, ON
CA$64,000 - CA$106,000Hybrid

About The Position

The Sr. Specialist, SRE – Compute Platforms serves as the enterprise technical owner for CTC's compute platforms, including IBM Mainframe (z/OS), AIX, IBM iSeries, IBM NOI, New Relic, and associated business-critical services and applications. This role is accountable for platform governance, service ownership, lifecycle strategy, observability governance, vendor oversight, and reliability outcomes across the compute environment. Acting as the internal authority for compute platforms, the position ensures services remain reliable, resilient, secure, observable, supportable, and aligned with enterprise technology strategy, risk management, and modernization objectives. Operational execution, administration, monitoring response, maintenance activities, and day-to-day infrastructure support are performed by HCL/Harmony and other designated support providers. The Sr. Specialist provides governance, strategic direction, technical leadership, and vendor accountability to ensure services are delivered in accordance with enterprise standards, contractual commitments, and business requirements. The role leads platform roadmap development, technology currency initiatives, service reliability improvements, observability strategy, operational governance, vendor management, and continuous improvement programs while promoting Site Reliability Engineering (SRE) principles, operational excellence, automation, and platform sustainability. The position drives platform maturity, reduces consultant and key-person dependency, and establishes clear ownership and accountability across the enterprise compute environment.

Requirements

  • 10+ years of experience in platform ownership, infrastructure governance, Site Reliability Engineering (SRE), service management, enterprise technology leadership, or related disciplines.
  • Strong background in Mainframe platform governance, including z/OS environments, lifecycle planning, capacity management, service ownership, and vendor-managed support models.
  • Experience governing enterprise monitoring, observability, event management, or operational intelligence practices across large-scale technology environments.
  • Strong understanding of service reliability principles, alert management, event correlation, service health monitoring, operational reporting, and risk management.
  • Knowledge of AIX, IBM iSeries, enterprise middleware, and application hosting environments supporting business-critical workloads.
  • Experience governing enterprise storage and backup services, including capacity planning, recovery validation, resiliency requirements, and restoration testing oversight.
  • Experience supporting disaster recovery governance, DR exercises, recovery reporting, remediation tracking, and resiliency planning.
  • Experience working within MSP-governed or vendor-managed service environments.
  • Strong understanding of incident, problem, change, lifecycle, operational risk, vendor governance, and service management processes.
  • Demonstrated ability to lead technical escalations, coordinate across multiple stakeholder groups, and hold vendors accountable to service expectations.
  • Strong communication skills with the ability to translate technical risk into clear operational and business impact.
  • Proven ability to influence technical direction, challenge operational assumptions, and drive accountability across diverse stakeholder groups.

Nice To Haves

  • Experience supporting IBM Mainframe environments, including z/OS, middleware, enterprise schedulers, file transfer platforms, or related infrastructure services.
  • Experience with observability and monitoring platforms such as IBM OMEGAMON, ITM/Tivoli, Netcool, New Relic, Dynatrace, Splunk, Elastic, ThousandEyes, or equivalent technologies.
  • Experience defining monitoring standards, observability strategies, alert rationalization programs, service health reporting, event correlation approaches, and operational intelligence improvements.
  • Experience supporting infrastructure audit controls, disaster recovery governance, compliance remediation activities, and evidence management processes.
  • Experience supporting highly regulated, high-availability, or business-critical technology environments.
  • Familiarity with ServiceNow-based intake, incident, event, problem, change, risk, vendor, and lifecycle management workflows.
  • Experience governing infrastructure refreshes, hardware lifecycle initiatives, maintenance planning, and data centre change activities.
  • Experience working with outsourced infrastructure providers, including SOW governance, SLA management, contract oversight, and vendor performance management.
  • Experience supporting SRE, service reliability, observability, platform governance, or operational excellence programs across hybrid infrastructure environments spanning Mainframe and distributed platforms.

Responsibilities

  • Provide governance and oversight of HCL/Harmony-delivered Mainframe services, including z/OS lifecycle planning, platform currency, capacity strategy, LPAR governance, and infrastructure roadmap planning.
  • Assess current platform, operating system, middleware, and Mainframe-supported application versions; identify lifecycle, supportability, operational, security, and compliance risks; and develop renewal and modernization strategies.
  • Lead platform lifecycle governance activities, including hardware refresh planning, operating system upgrade strategy, maintenance governance, change oversight, and long-term roadmap alignment.
  • Govern service ownership, support models, operational accountability, and vendor-delivered support services for Mainframe-hosted and Mainframe-dependent applications.
  • Define, govern, and periodically review monitoring and observability requirements for Mainframe infrastructure, middleware, and business-critical services.
  • Provide governance and oversight of disaster recovery readiness, resiliency planning, recovery testing, and recovery capability validation delivered by managed service providers.
  • Identify and drive resolution of operational ownership, support model, documentation, and Statement of Work (SOW) gaps through the appropriate vendors, support teams, and stakeholders.
  • Provide technical leadership and governance during major incidents, problem investigations, and corrective action planning, ensuring responsible teams execute required remediation activities.
  • Serve as the primary technical escalation and governance authority for Mainframe-related risks, service concerns, lifecycle issues, and vendor performance matters.
  • Provide governance and strategic oversight of enterprise monitoring, event management, and observability capabilities.
  • Establish governance standards for monitoring coverage, alerting requirements, event correlation, escalation models, service health dashboards, and operational reporting.
  • Drive continuous improvement of observability maturity, service visibility, monitoring effectiveness, synthetic monitoring capabilities, and operational intelligence.
  • Govern event management practices, including alert quality, escalation effectiveness, incident correlation, operational readiness, and service monitoring standards.
  • Define strategic direction and adoption roadmaps for observability platforms, monitoring technologies, event management tooling, and automation capabilities.
  • Act as the Compute SRE representative for enterprise monitoring strategy, operational intelligence, and observability initiatives.
  • Provide technical governance, vendor oversight, and escalation leadership for HCL/Harmony-managed services across Mainframe, AIX, IBM iSeries, storage, backup, monitoring, and supporting infrastructure platforms.
  • Validate vendor-delivered services against contractual obligations, SOW commitments, service level expectations, operational controls, monitoring standards, and governance requirements.
  • Lead vendor performance reviews, operational scorecards, service reporting reviews, incident follow-ups, and continuous improvement initiatives.
  • Identify ownership gaps, operational risks, monitoring deficiencies, contractual concerns, and escalation requirements for leadership review.
  • Ensure vendor-delivered services are properly documented, transitioned, supportable, and aligned with the enterprise operating model.
  • Drive vendor accountability for service reliability, incident response effectiveness, operational reporting quality, remediation commitments, and lifecycle objectives.
  • Ensure operational documentation, runbooks, support procedures, monitoring standards, and technical documentation are established, maintained, and periodically reviewed by responsible support teams and service providers.
  • Support audit, compliance, and control activities related to backup, recovery, disaster recovery, monitoring controls, operational governance, and infrastructure lifecycle management.
  • Identify opportunities to improve reliability, observability, automation, resiliency, operational efficiency, and platform sustainability.
  • Participate in change, incident, problem, risk, lifecycle, and service governance forums as the Compute platform owner.
  • Provide reporting and recommendations on platform risks, technology currency gaps, vendor performance, monitoring effectiveness, remediation priorities, lifecycle initiatives, and service readiness.
  • Lead continuous improvement initiatives focused on service maturity, reliability, observability, sustainability, risk reduction, and operational excellence.

Benefits

  • Comprehensive benefits and retirement programs
  • Performance incentives
  • Continuing Education Programs
  • Other perks to support your well-being
  • Career growth opportunities
  • Product discounts
  • Mental health benefits in the amount of $5,000 per year for benefits-eligible employees and their families
  • Total well-being, and mental health tools and resources for all employees
  • Store discounts
  • Supported learning through our Triangle Learning Academy
  • Canadian Tire Profit Sharing
  • Retirement and savings programs for eligible employees
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service