Site Reliability Engineer - Application Support

CapgeminiToronto, ON
CA$80,000 - CA$90,000Onsite

About The Position

This role is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability, availability, and performance of enterprise applications running across Windows, UNIX/Linux, OpenShift, and PostgreSQL environments. Capgemini means choosing a company where you will be empowered to shape your career, supported and inspired by a collaborative community, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world.

Requirements

  • Strong hands-on experience supporting and troubleshooting: Windows Server
  • Strong hands-on experience supporting and troubleshooting: UNIX/Linux environments
  • Experience with OpenShift Container Platform (OCP), including: Application deployment and support
  • Experience with OpenShift Container Platform (OCP), including: Pod and container troubleshooting
  • Experience with OpenShift Container Platform (OCP), including: Resource monitoring and performance analysis
  • Experience with OpenShift Container Platform (OCP), including: OpenShift administration fundamentals
  • Strong experience with PostgreSQL, including: Database monitoring and support
  • Strong experience with PostgreSQL, including: SQL query analysis and troubleshooting
  • Strong experience with PostgreSQL, including: Performance tuning and optimization
  • Strong experience with PostgreSQL, including: Backup and recovery validation
  • Strong experience with PostgreSQL, including: Database connectivity troubleshooting
  • Experience in Application Support, Production Support, SRE, or Platform Operations roles.
  • Strong understanding of incident, problem, and change management processes.
  • Experience supporting mission-critical applications in large enterprise environments.
  • Ability to analyze logs, alerts, metrics, and system trends to rapidly resolve issues and improve service reliability.
  • Knowledge of monitoring and observability tools for proactive system management.
  • Strong analytical, troubleshooting, and communication skills.

Nice To Haves

  • Experience in banking, financial services, or other highly regulated industries.
  • Exposure to automation and scripting using PowerShell, Bash, Python, or Shell Scripting.
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of cloud-native technologies and hybrid cloud environments.
  • Experience with monitoring tools such as Splunk, Dynatrace, AppDynamics, Grafana, Prometheus, or similar platforms.
  • Understanding of ITIL processes, SLOs, SLIs, and reliability engineering principles.

Responsibilities

  • Provide Level 2/Level 3 application and infrastructure support for enterprise applications hosted on Windows, UNIX/Linux, OpenShift, and PostgreSQL platforms.
  • Monitor application health, server performance, database availability, and OpenShift container workloads to ensure optimal system reliability and uptime.
  • Troubleshoot and resolve production incidents across operating systems, middleware, containers, and databases.
  • Perform root cause analysis (RCA) and implement preventative measures to reduce recurring incidents.
  • Support application deployments, environment management, and release activities across multiple environments.
  • Manage and support OpenShift platform operations, including container deployments, pod troubleshooting, scaling, and monitoring.
  • Monitor and support PostgreSQL databases, including performance tuning, query analysis, backup validation, and connectivity troubleshooting.
  • Collaborate with development, infrastructure, database, and cloud teams to drive operational excellence.
  • Develop and maintain automation scripts, operational runbooks, dashboards, and support documentation.
  • Participate in on-call support rotations and major incident management activities.
  • Identify and implement opportunities for process improvement, automation, and operational efficiency.

Benefits

  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave
  • Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
  • Other benefits as provided by local policy and eligibility
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service