Director, AI Platform Engineering

BMOToronto, ON
CA$121,600 - CA$211,800

About The Position

BMO is building a dedicated AI Engineering function to deliver the platform capabilities that make enterprise AI safe, governed, and scalable across its business domains and regulatory regimes. This role involves leading a team that designs, ships, and operates the core infrastructure for AI governance and enforcement, including the AI Gateway, Policy Engine, Identity Fabric, AI Registry, Guardrails Runtime, and AI Observability. The focus is on the governed platform AI workloads run on, and the runtime evidence proving policy adherence, rather than the AI models or applications themselves. The ideal candidate is a hands-on technical leader with experience building and operating scalable platform capabilities, emphasizing operability and regulatory defensibility. They will leverage existing assets like a developer portal, AI registry, policy-as-code, and gateway integrations to formalize and scale them into an enterprise-grade platform, partnering across Security, Architecture, DevOps, and domain teams.

Requirements

  • 8+ years in technical platform, infrastructure, or AI/ML engineering roles in a large enterprise.
  • 4+ years leading and managing engineering teams.
  • Proven organizational leadership: building and scaling engineering teams, workforce planning, hiring, succession planning, and structuring squads.
  • Demonstrated team building across blended teams (net-new hires with internal engineers).
  • Strong mentoring and coaching track record.
  • Ability to establish and sustain a healthy, inclusive team culture aligned to BMO Values.
  • Experience leading through change and ambiguity.
  • Conflict resolution and cross-team influence.
  • Demonstrated experience building and operating platform capabilities at scale (API gateways, policy/authorization systems, identity/workload-identity infrastructure, observability pipelines, or equivalent shared services).
  • Strong knowledge of GenAI platform engineering (LLM/AI gateways, model routing and abstraction, RAG and agentic patterns, guardrails, AI evaluation approaches).
  • Hands-on experience with policy-as-code and authorization systems (Cedar, OPA/Rego, or equivalent) and GitOps-based distribution.
  • Experience with workload identity and zero-trust patterns (SPIFFE/SPIRE, mTLS, token exchange, federated identity) or strong adjacent identity/security engineering depth.
  • Strong observability engineering background (OpenTelemetry, distributed tracing, telemetry pipelines).
  • Multi-cloud fluency (AWS and Azure preferred), cloud-native architecture, containerization/Kubernetes, and Infrastructure as Code.
  • Hands-on familiarity with modern AI/ML tooling (e.g., Bedrock, Azure OpenAI, SageMaker, Databricks, MLflow, LangChain, or equivalents).
  • Proven CI/CD, DevSecOps, and MLOps/LLMOps delivery experience.
  • Solid grounding in Responsible AI, AI/data governance, privacy, and ideally model-risk management and financial-services regulatory expectations.
  • Executive-grade communication and relationship management across technical and senior-leadership audiences.
  • Strategic and organizational management skills, including multi-year roadmap planning, budgeting, forecasting, and vendor engagement.
  • Critical thinker with strong analytical, problem-solving, and prioritization abilities.
  • Bachelor's degree in Computer Science, Software Engineering, or a related technical discipline.

Nice To Haves

  • Master's degree preferred.
  • Relevant certifications: cloud (AWS/Azure/GCP) architecture or ML/AI certifications, Kubernetes (CKA/CKAD), security/identity certifications, or enterprise architecture (TOGAF or equivalent).

Responsibilities

  • Productionize the developer portal and deliver a federated AI Registry spanning agents, models, tools, channels, and evaluations, with self-service onboarding and lifecycle workflows.
  • Develop policy-as-code infrastructure (Cedar/OPA), a policy compilation and GitOps distribution pipeline, risk-tiered approval workflows, and a policy simulation environment.
  • Establish a multi-pipeline architecture for operational, security, and compliance telemetry, including OpenTelemetry GenAI conventions, cross-pipeline trace correlation, and a tamper-evident audit lake for regulator-ready evidence.
  • Implement certification workflows, automated compliance scoring, decommission governance, and evidence generation for architecture and model-risk review.
  • Deploy domain-hub gateway runtime across multiple clouds, including an inline enforcement engine with request-time policy evaluation, routing, residency, budget/quota controls, and circuit breakers, operating within strict latency budgets.
  • Develop a multi-stage safety pipeline (input moderation, prompt-injection defense, PII handling, output validation, hallucination detection, policy enforcement) with bilingual (EN/FR) parity and behavioral guardrails for agentic workloads.
  • Implement workload identity for AI (SPIFFE/SPIRE), token-exchange bridging, per-domain trust boundaries, enterprise identity integration, and cross-cloud token federation with zero-trust attestation.
  • Deliver a production-hardened Developer Portal and federated AI Registry with sub-5-day self-service onboarding within the first 12 months.
  • Operationalize an AI Gateway in a selected business domain, meeting tiered latency targets within the first 12 months.
  • Implement policy-as-code infrastructure distributing domain-scoped policy bundles via GitOps, with a working simulation sandbox within the first 12 months.
  • Establish a runtime evidence pipeline producing lineage-stamped, audit-ready traces aligned to model-risk and regulatory expectations within the first 12 months.
  • Scale the team from an initial core (8–12 FTE) toward steady-state through hiring and reallocation of internal engineers within the first 12 months.
  • Operate what the team builds, designing for operability and engineering support from the start.
  • Enable domains to consume enforcement infrastructure while domains own their workloads.
  • Produce regulatory evidence at runtime through instrumented infrastructure.
  • Ensure teams own outcomes (e.g., 'Identity Fabric works across all platforms') rather than specific technologies.

Benefits

  • health insurance
  • tuition reimbursement
  • accident and life insurance
  • retirement savings plans
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service