Lead DevOps & Platform Engineer

Inspired HRVancouver, BC
CA$140,000 - CA$160,000Hybrid

About The Position

We are a growing Vancouver technology startup building proprietary, AI-enabled software and smart digital solutions. We use AI extensively across product development and engineering, while protecting our source code, confidential information, customer data, brands, and other intellectual property. We are now investing in the cloud platform, delivery systems, and operational discipline required to turn rapid innovation into secure, reliable products. As our first Lead DevOps & Platform Engineer, you will establish and shape the technical foundation of our software development lifecycle — owning the evolution of our Azure platform and software-delivery foundation. This is a hands-on role for someone who can move comfortably between architecture, implementation, incident response, security, automation, and technical leadership. You will inherit Azure-hosted products supported by six international developers and a flexible group of Canada-based contractors. Core capabilities, including Infrastructure as Code, centralized observability, mature GitHub-based CI/CD, and a sustainable production-support model, still need to be established. You will assess the current state, choose pragmatic tools and practices, and bring the process discipline needed to create a foundation that fosters high-quality digital products. You'll work closely with our Product Manager and development team to implement the standards, tooling, and processes that will scale as our product offerings grow. There are no direct reports initially, but you will set the technical standard for DevOps — providing technical direction across the delivery team and determining how work is assigned and executed once it's been prioritized by the Product Manager — while helping leadership determine future staffing needs. As the company grows, you will participate in hiring and onboarding additional engineers and may have the opportunity to move into formal people leadership.

Requirements

  • Substantial hands-on experience, typically seven or more years, in DevOps, platform engineering, Full Stack Development, UI/UX design processes, Scrum facilitation, site reliability engineering, cloud infrastructure, and creation of a variety of digital products and service or business-critical systems.
  • Comprehensive knowledge of Agile/SAFe Agile principles and best practices used to manage software development lifecycles and improve high quality digital products and services.
  • Demonstrated ownership of production workloads in Microsoft Azure, with sound judgment across architecture, networking, identity, security, resilience, performance, and cost.
  • Deep experience building GitHub Actions pipelines and administering GitHub repository and organization governance.
  • Production experience with Terraform and/or Bicep and the engineering practices needed to keep Infrastructure as Code safe, reviewable, reusable, and maintainable.
  • Experience establishing monitoring, logging, tracing, alerting, incident response, backup, and disaster-recovery practices in an early-stage or immature environment.
  • Strong cloud-security fundamentals, including Entra-based identity and access, secrets management, network controls, vulnerability management, and secure software-supply-chain practices.
  • Professional experience supporting AI- or ML-enabled applications in production and using AI-assisted engineering tools or agents responsibly, including their security, privacy, observability, evaluation, reliability, and cost considerations.
  • Ability to automate operational work using PowerShell, Bash, Python, or another suitable language.
  • Demonstrated technical leadership across distributed teams with an ability to build positive working relationships with contractors, or external partners, and clear written and verbal communication skills suitable for technical and non-technical audiences.
  • A degree, diploma, or certificate in computer science, engineering, or a related discipline, or an equivalent combination of education and relevant professional experience.

Nice To Haves

  • Azure-hosted AI services, managed model APIs, retrieval or search components, AI gateways, model or prompt evaluation, and LLMOps/MLOps practices.
  • Azure Monitor, Application Insights, Log Analytics, OpenTelemetry, Grafana, Datadog, or comparable observability technologies.
  • Building or scaling SaaS/PaaS products in a startup or growth-stage company.
  • Healthcare, pharmacy, or another regulated environment, including implementation of technical controls and evidence supporting PIPEDA, HIPAA, SOC 2 readiness, or comparable obligations.
  • Docker and Kubernetes or AKS where orchestration is justified; Azure, cloud-security, infrastructure, FinOps, or related certifications are also valued.
  • Comprehensive experience working in an ITIL 4 IT Service Management environment.

Responsibilities

  • Implement industry-standard DevOps & Agile/SAFe Agile standards, tools, and governance frameworks, and lead engineering teams in applying these best practices to the software development lifecycle to meet business requirements and user needs.
  • Own the architecture, reliability, security, performance, and cost management of the Microsoft Azure platform across development, test, and production environments.
  • Implement repeatable Infrastructure as Code using Terraform and/or Bicep, including reviewed modules, environment separation, secure state, drift awareness, and controlled change practices.
  • Standardize GitHub Actions pipelines, repository governance, branch protections, reusable workflows, release controls, and reliable rollback paths.
  • Embed automated testing, code-quality checks, web application vulnerability scanning, monitoring and remediation of CVEs/zero-day threats, secrets protection, and deployment safeguards into the delivery lifecycle.
  • Select and implement an observability stack from the ground up, covering logs, metrics, traces, dashboards, actionable alerts, service ownership, and appropriate service-level objectives.
  • Using Azure and Entra ID capabilities, establish and support secure cloud identity and access management policies including secrets, network-security, conditional access policies, role-based access, privileged identity management (PIM), and Azure Key Vault.
  • Lead production incident response, blameless reviews, runbooks, and continual reliability improvements. Critical incidents have historically been infrequent; initially, this role will be a primary after-hours technical escalation point and will establish a more sustainable future coverage model.
  • Define and test recovery objectives, disaster-recovery procedures, and business-continuity plans, while improving Azure cost visibility, budgets, alerts, and optimization.
  • Contribute to the implementation and management of ITSM best practices using Jira Service Management for Change, Incidents, Problem, and Service Requests.
  • Establish secure, approved AI-assisted development workflows across coding, testing, review, documentation, infrastructure work, troubleshooting, and incident analysis, with human validation for consequential changes.
  • Create practical guardrails that protect proprietary source code, credentials, confidential information, customer data, and regulated data when engineers use AI tools and agents.
  • Partner with product and development teams to productionize AI-enabled services on Azure, including repeatable environments, identity, secrets, evaluation and release gates, versioning, rollback, monitoring, rate limits, consumption, and cost controls.
  • Evaluate emerging AI platform and engineering tools based on measurable business value, security, reliability, and cost rather than novelty, and use AI responsibly to reduce operational toil.
  • Translate product priorities into architecture, technical plans, sequencing, estimates, and delivery standards. Product leadership owns business outcomes and priority; this role owns the technical approach and operational readiness.
  • Coordinate technical implementation and execution of work by the broader development team comprised of six international developers and rotating Canada-based contractors through architecture and code reviews, mentoring, documentation, and clear engineering standards.
  • Maintain various documentation; architecture diagrams, ITSM records, runbooks, operating procedures, executive summaries of platform risks & investment requirements, progress reports, communications, and presentations.
  • Assess capability and capacity gaps, recommend the next-stage platform organization, and participate in selecting, interviewing, and onboarding future hires.

Benefits

  • Eligibility for a performance-based bonus.
  • Extended health and dental coverage.
  • Paid vacation and statutory holidays in accordance with company policy.
  • Employer-supported training or certifications where aligned with the role and approved business needs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service