DevOps Engineer - BTP Site Reliability Engineering team

SAPMontreal, QC
CA$97,800 - CA$166,200Hybrid

About The Position

We are looking for an engineer to join an already established SRE team for the SAP Business AI Platform. As a Site Reliability Engineer, you will have the opportunity to operate and support business critical Cloud services. As part of your daily job, you will proactively monitor the service behavior and identify areas for improvement. You will participate in the development of tools for monitoring and troubleshooting cloud services built on latest open source and SAP technologies, following SRE principles.

Requirements

  • Experience with Kubernetes and good understanding of container technologies.
  • Understanding of modern cloud architectures (experience with Cloud Platforms such as AWS, Azure, GCP are a plus).
  • Experience with Unix/Linux operating system
  • Scripting skills, CI/CD (ArgoCD, Concourse, Github Actions and are a plus) - enthusiasm for automation - make the computers do the work for you.
  • Experience using AI-assisted engineering tools (e.g., Claude Code CLI, GitHub Copilot, or similar) to improve troubleshooting, automation, documentation, root cause analysis, and operational efficiency.
  • 2+ years experience in SRE.
  • Working efficiently in emergency situations. Affinity to quickly analyze and solve problems in a global team setup.
  • Excellent team player, passionate about his/her work, self-motivated and driven.
  • Excellent communication skills - precise, based on facts.
  • Fluency in English.

Nice To Haves

  • Coding experience with Python, GO, Bash
  • CKA/CKAD/CKS certifications
  • Experience with modern monitoring, logging, and alerting tools (Grafana, Prometheus, Kibana, Loki, Splunk On-Call, Dynatrace)
  • Security best practices for application development and operations in a public Cloud Environment
  • Contribution to open-source projects

Responsibilities

  • Act as technical expert during Live site incidents (downtimes of supported services in scope), investigate and solve incidents on a deep technical level.
  • Drive root cause analysis and follow-up improvements to prevent issues from reoccurring.
  • Perform in-depth troubleshooting and log analysis to identify and solve complex issues in accordance with internal and external SLAs.
  • Build software-based solutions to address improvements in service reliability and stability.
  • Enhance infrastructure and platform monitoring by gathering system metrics (4 Golden Signals) and implementing tools for recovery.
  • Integrate and collaborate closely with development teams and work with them on outputs from Postmortems and product improvements.
  • Learn new technologies and keep up to date with latest development increments.
  • Create and maintain technical documentation.
  • Define, advocate, apply SRE best practices.
  • Participate in the on-call rotation (follow the sun approach) to react to major incidents. On-call has a special compensation package.

Benefits

  • Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
  • On-call has a special compensation package.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service