Principal Engineer, AI Platform

BMOToronto, ON
CA$103,200 - CA$192,000

About The Position

BMO is building the platform capabilities that make enterprise AI safe, governed, and scalable. We are seeking experienced Principal/Senior engineers to build and operate the core infrastructure that governs how AI runs at BMO: the AI Gateway, Policy Engine, Identity Fabric, AI Registry, Guardrails Runtime, and AI Observability. This is a build-and-run engineering role. You will own capabilities end to end; designing, shipping, and operating them in production, including on-call. You will not build the AI models or applications themselves (those are domain-owned); you build the governed platform they run on and the runtime evidence that proves they run within policy, across AWS, Azure, and Microsoft AI surfaces, under OSFI and OCC expectations. You are a hands-on engineer who has built shared platform services at scale, cares deeply about operability, latency, and correctness, and understands that in a regulated bank the infrastructure must produce its own evidence. You are energized by taking real engineering assets that includes an existing developer portal, an AI registry, a body of policy-as-code, and gateway integrations, and hardening, scaling, and governing them into enterprise-grade platform capabilities. You raise the technical bar for those around you and mentor as you build.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, or a related technical discipline (Master's preferred).
  • 8+ years of software/platform engineering experience (Principal), or 5+ years (Senior), with substantial time building and operating shared platform services at enterprise scale.
  • Demonstrated experience operating production infrastructure with real SLOs and on-call ownership, ideally in a regulated industry (financial services strongly preferred).
  • Depth in one or more of: API gateways / traffic enforcement; policy-as-code and authorization; workload identity / zero-trust; observability and telemetry pipelines; audit/compliance data platforms.
  • Strong distributed-systems and platform-engineering fundamentals: latency-sensitive request paths, resilience patterns (circuit breakers, failover), multi-tenancy, and high availability.
  • Strong programming skills (Python and/or Go preferred; TypeScript/Java an asset) for building performant services, APIs, and integrations.
  • Cloud-native architecture across AWS and Azure: containers/Kubernetes, service mesh, and Infrastructure as Code (CDK, Terraform, CloudFormation/ARM).
  • Robust CI/CD, GitOps, and DevSecOps practice; Git-based workflows (Bitbucket/GitHub), Jira, Confluence.
  • Working knowledge of GenAI platform patterns: LLM/AI gateways, RAG and agentic patterns, foundation models, embeddings, and guardrails — sufficient to build the infrastructure they depend on.
  • Familiarity with AI/ML platforms (Bedrock, Azure OpenAI, SageMaker, MLflow) and orchestration frameworks (LangChain, LlamaIndex).
  • Grounding in Responsible AI, AI/data governance, privacy, cloud security, and IAM as applied to AI workloads.
  • Strong communication and collaboration across engineering, security, architecture, and domain teams.
  • A critical thinker with strong analytical and problem-solving skills.
  • Self-directed, comfortable with ambiguity and a fast-evolving mandate.
  • Able to deliver complex work under tight timelines; participates in on-call rotation for owned services.

Nice To Haves

  • Master's degree
  • AWS Certified Solutions Architect (Associate/Professional) / ML – Specialty
  • Microsoft Certified: Azure Solutions Architect Expert / Azure AI Engineer Associate
  • Kubernetes (CKA/CKAD); HashiCorp Terraform Associate
  • Security/identity certifications (relevant to Identity Fabric roles)

Responsibilities

  • Own capabilities end to end — design, implement, test, ship, and operate production infrastructure, including on-call ownership of what you build (no separate run team).
  • Engineer for operability and defensibility from day one — instrumentation, SLOs, latency budgets, failure modes, and runtime evidence built in, not bolted on.
  • Build the APIs, SD’able interfaces, and integrations through which domains, DevOps pipelines, and enterprise systems consume platform capabilities.
  • Implement policy enforcement, guardrails, identity attestation, and audit as first-class engineering concerns — correct, performant, and provable.
  • Ensure every capability produces runtime evidence connecting AI activity to policy enforcement, identity, and lineage for model-risk and regulatory review (OSFI E-23, OCC).
  • Assess emerging AI infrastructure, foundation-model access patterns, and standards; make deliberate, cost-aware engineering choices.
  • Mentor and raise the bar — set engineering standards, review designs and code, and grow depth across the team.
  • Partner closely with AI Developer Experience (so domains can consume what you build), AI Security and AI SDLC (embedded specializations), and the Senior AI Architect (architectural coherence).

Benefits

  • health insurance
  • tuition reimbursement
  • accident and life insurance
  • retirement savings plans
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service