Senior Lead Engineer, Platform Operations & Observability

McKessonMontreal, QC
CA$94,400 - CA$157,300Hybrid

About The Position

McKesson is seeking a Senior Lead Engineer, Platform Operations & Observability to lead the reliability, observability, and operational excellence of enterprise healthcare technology platforms. In this role, in accordance with Application Monitoring & Observability Lead, you will drive monitoring strategies, incident management practices, change governance, and root cause analysis initiatives while helping build scalable, secure, and resilient systems. You will collaborate with software engineering, platform, security, and operations teams to improve service reliability, automate operational processes, and establish best practices for production readiness. This position also provides technical leadership and mentorship to engineering teams while influencing reliability standards and long-term platform strategy.

Requirements

  • 7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • Experience supporting large-scale production environments and enterprise applications.
  • Hands-on experience with monitoring, observability, logging, alerting, and application performance monitoring tools.
  • Proven experience leading incident management and production support activities.
  • Experience performing root cause analysis and implementing preventive solutions.
  • Experience with CI/CD, automation, DevOps practices, and software delivery pipelines.
  • Experience with microservices, APIs, distributed systems, and cloud-based architectures.
  • Lead production readiness reviews and operational acceptance activities prior to major releases.

Nice To Haves

  • Experience with tools such as Dynatrace, Prometheus, Dotcom Monitor or similar observability platforms.
  • Experience with Kubernetes, containers, and cloud platforms such as Azure.
  • Knowledge of ITIL-aligned incidents, problems, and change management practices.
  • Experience defining SLAs, MTTR and service reliability metrics.
  • Demonstrated technical leadership and mentorship of engineering teams.
  • Experience operating in regulated or highly compliant environments.
  • Experience driving platform modernization and operational excellence initiatives.
  • Strong analytical, troubleshooting, continuous improvement and stakeholder communication skills.
  • Fair understanding and mastery of Ai tools (Copilot, Rovo) and Ai agents.

Responsibilities

  • Lead monitoring and observability strategies across enterprise applications, platforms, and services.
  • Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions.
  • Drive incident management processes, major incident response, escalation coordination, service restoration activities and postmortem incident.
  • Conduct root cause analysis (RCA) investigations and lead corrective and preventive action planning.
  • Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness.
  • Lead change management reviews and promote safe deployment and release practices.
  • Ability to execute regression and validation test plans following each production deployment.
  • Provide technical leadership, coaching, and mentoring to engineers while establishing engineering best practices.
  • Influence architecture, automation, CI/CD, and operational excellence initiatives supporting enterprise platforms.

Benefits

  • competitive compensation package
  • annual bonus
  • long-term incentive opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service