SRE Associate (Hybrid)

Morgan StanleyMontreal, QC
Hybrid

About The Position

We’re seeking someone to join our team as a Site Reliability Engineering Specialist to support the reliability, performance, and operational stability of business-critical applications, while helping drive automation and the adoption and operation of AI-enabled technologies and modern engineering practices. In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Software Production Management & Reliability Engineering II position at Associate, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime. Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.

Requirements

  • At least 2 years' relevant experience would generally be expected to find the skills required for this role.
  • Strong knowledge of Linux/Unix operating systems and production support environments.
  • Strong scripting and automation skills using Python, Bash, Shell scripting, or similar technologies.
  • Understanding of networking concepts, protocols, APIs, REST services, system integrations, and troubleshooting methodologies.
  • Experience with monitoring, observability, and reliability engineering concepts and tools such as OpenTelemetry, Grafana, and Loki.
  • Experience using modern integrated development environments (IDEs) and version control platforms, including Visual Studio Code and Git.
  • Familiarity with containerization technologies such as Docker and Kubernetes, and exposure to cloud technologies and automation frameworks.
  • Basic understanding of Generative AI concepts, including Large Language Models (LLMs), AI agents, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Familiarity with AI-assisted development tools such as GitHub Copilot, Microsoft Copilot, or similar productivity platforms, with an interest in applying AI and automation to improve operational efficiency.
  • Knowledge of French and English is required.

Responsibilities

  • Work independently to solve ambiguous operational and technical problems while supporting business-critical applications.
  • Collaborate with development, QA, product management, and business stakeholders to support applications throughout their lifecycle.
  • Escalate and manage major incidents, perform root cause analysis, and implement corrective actions to improve reliability.
  • Develop and maintain automation, scripts, operational tooling, documentation, and runbooks to improve efficiency and reduce manual effort.
  • Perform platform lifecycle activities, including upgrades, patching, archival, cleanup, and ongoing maintenance.
  • Support the evaluation, deployment, and operation of third-party software, AI-enabled platforms, and emerging technologies.
  • Partner with global teams to improve support processes, operational effectiveness, and continuous improvement initiatives.

Benefits

  • Comprehensive employee benefits and perks in the industry.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service