Senior Platform Engineer

Howden•Charlotte, NC
•Remote

About The Position

The Senior Platform Engineer builds and runs the Azure platform that Howden's US AI and data workloads sit on. This is a hands-on engineering role: you will design, code, and operate the landing zones, the data platform, the AI runtime services, and the networking, identity, and security controls that let engineering teams ship AI capability safely and quickly. The scope spans three connected layers. The Azure foundation: subscriptions, landing zones, networking, private connectivity, identity, and policy, all expressed as Terraform. The data platform: Databricks, Lakehouse storage, Unity Catalog, ingestion and orchestration, and the governed data products that AI systems consume. The AI platform: Azure AI Foundry, the AI Gateway, agent runtime and hosting, model deployment and quota management, and the control surfaces such as Microsoft Agent 365 and the agent registry that make agent estates visible and governable. You will work from architecture direction set by the Leads, Group Architecture, and InfoSec, and you will own the implementation end to end. Expect to spend most of your time in Terraform, in pipelines, and in Azure, with a heavy emphasis on agent-assisted development: using coding agents and AI development tooling to move faster than a conventional infrastructure pace and knowing when to trust the output and when to rewrite it. This role is self-directed. You will be handed an outcome and a rough shape, and you are expected to break it down, sequence it, unblock yourself, and deliver production-quality work collaborating with other team members.

Requirements

  • Bachelor's degree in computer science, Engineering, Information Systems, or a related technical field, or equivalent hands-on experience.
  • 6+ years of professional experience in cloud platform, infrastructure, or data engineering, including 3+ years building and operating production workloads on Azure.
  • Deep hands-on Azure experience across compute, storage, networking, identity, and integration services, with production ownership rather than project exposure.
  • Strong Terraform experience provisioning Azure infrastructure at scale: modules, state management, multi-environment workflows, versioning, and drift management.
  • Hands-on experience designing and implementing Azure networking and connectivity: hub-and-spoke, VNets, private endpoints and Private Link, DNS, firewalls, and NSGs.
  • Hands-on Azure security and identity experience: Entra ID, managed identities and service principals, RBAC design, Key Vault, encryption, and secure connectivity patterns.
  • Production experience with Azure data platform tooling, including Databricks and at least one of Data Factory, Microsoft Fabric, or Synapse, covering ingestion, transformation, orchestration, and cost management.
  • Working experience with lakehouse and Delta patterns, data modelling, and strong SQL, plus hands-on work with both structured and unstructured enterprise data.
  • Production experience with Azure AI services such as Azure AI Foundry, Azure OpenAI Service, or Azure AI Search, including deployment, quota, and access management.
  • Strong scripting and automation ability in Python, plus PowerShell or Bash, and comfort reading and contributing to application code.
  • Experience building and maintaining CI/CD pipelines in Azure DevOps, GitHub Actions, or GitLab, including approval gates and automated testing.
  • Working experience with containers and cloud-native deployment patterns: Docker, serverless functions, and managed container services such as Container Apps or AKS.
  • Experience instrumenting and supporting production systems with Azure Monitor, Application Insights, or comparable observability tooling, and owning cost visibility for the platform you run.
  • Solid engineering fundamentals: Git workflow, code review, automated testing, debugging, and writing code other engineers can maintain. Expect to walk through the code you have written.
  • Practical daily use of AI coding assistants or coding agents in a professional engineering workflow.
  • Demonstrated ability to take an outcome and deliver it independently, breaking down the work and resolving blockers without close direction.
  • Experience contributing to architecture and technical trade-off discussions, presenting options and recommendations to architects, security, or engineering leadership.
  • Clear written and verbal communication with technical and non-technical colleagues.

Nice To Haves

  • Microsoft Azure certifications (Solutions Architect Expert, Azure Administrator, DevOps Engineer Expert, AI Engineer Associate, or Data Engineer Associate) or comparable cloud credentials.
  • Databricks certification and experience with Unity Catalog, cluster policies, and workspace governance at enterprise scale.
  • Hands-on experience with AI gateway patterns: Azure API Management AI Gateway, model routing, token-based throttling, semantic caching, and chargeback models.
  • Experience with Microsoft Agent 365, Entra Agent ID, or comparable agent identity, inventory, and lifecycle tooling.
  • Experience with agent runtimes and orchestration frameworks such as Azure AI Agent Service, Semantic Kernel, LangGraph, or AutoGen, and with Model Context Protocol (MCP) servers and tool integration.
  • Experience with data governance tooling such as Microsoft Purview, including classification, lineage, and access policy.
  • Experience with Azure landing zone frameworks, Azure Policy as code, and enterprise-scale governance in a multi-subscription estate.
  • Experience with Kubernetes at production scale, including networking, scaling, and workload identity.
  • Experience with ServiceNow or comparable ITSM and GRC integration for change, approval, and CMDB workflows.
  • Experience working to a defined SDLC or AI governance framework, and with regulatory regimes such as SOX, GDPR, NIST AI RMF, ISO/IEC 42001, or the EU AI Act.
  • Familiarity with Azure integration services: API Management, Logic Apps, Service Bus, Event Hubs.
  • Background in insurance, financial services, or another regulated industry.

Responsibilities

  • Design, build, and operate Azure landing zones for AI and data workloads: subscription and management group structure, naming and tagging standards, Azure Policy, RBAC models, and cost boundaries.
  • Provision and run the compute and hosting layer for AI services: Azure Container Apps, AKS, App Service, Azure Functions, and container registries, with sensible scaling, resilience, and resource configuration.
  • Build shared platform services that product teams consume: API Management, Key Vault, Service Bus, Event Hubs, Storage, Cosmos DB, and managed identity patterns.
  • Own capacity, quota, and region strategy for AI and data services, including data residency and data zone constraints.
  • Keep environments reproducible. Anything created by hand in the portal gets replaced by code.
  • Build and operate Azure AI Foundry as a shared platform capability: resource and project topology per environment and data zone, model deployments, quota and throughput management, content safety configuration, and connection management.
  • Implement and run the AI Gateway layer using Azure API Management AI Gateway and Foundry control plane capabilities: model routing, token-based rate limiting, semantic caching, cost attribution by team and workload, key and identity management, and usage telemetry.
  • Stand up and operate the runtime that hosts AI agents: container and serverless hosting, orchestration runtimes, tool and MCP server connectivity, secrets and credential handling, and environment promotion.
  • Implement agent identity and control surfaces including Microsoft Agent 365, Entra Agent ID, and the agent registry: registration, ownership, entitlement, lifecycle, and decommissioning of agents in the estate.
  • Integrate the platform with governance and service management tooling so that agent inventory, approvals, and change records stay current rather than being maintained by hand.
  • Build the guardrails that make self-service safe: landing-zone templates, workload onboarding automation, quotas, policy checks, and default observability wired in from day one.
  • Build and operate the Databricks platform: workspace topology, clusters and serverless compute, Unity Catalog, cluster policies, job orchestration, and cost controls.
  • Implement Lakehouse storage and medallion layering on ADLS Gen2 with Delta, including partitioning, retention, and lifecycle management.
  • Build ingestion and transformation pipelines from insurance source systems using Data Factory, Databricks workflows, Microsoft Fabric, or event-driven patterns, with lineage, versioning, change detection, and reconciliation.
  • Implement data governance in the platform: catalog and metadata management, classification, data quality checks, access controls, masking, and audit trails, working with Purview or equivalent tooling.
  • Expose governed data and retrieval layers that AI systems depend on, including vector stores and search indexes, and keep them current through automated refresh.
  • Write the SQL, Python, and transformations needed to answer your own data questions rather than waiting on another team.
  • Design and implement Azure network architecture for AI and data workloads: hub-and-spoke topology, VNets and subnets, private endpoints and Private Link, DNS, firewalls, NSGs, WAF, and egress control.
  • Implement identity and access end to end: Entra ID, managed identities, service principals, workload identity federation, OAuth 2.0 flows, conditional access, and least-privilege RBAC across Azure, Databricks, and AI services.
  • Secure secrets, keys, and certificates through Key Vault with rotation and automated distribution to workloads.
  • Implement data protection controls: encryption at rest and in transit, customer-managed keys where required, PII handling, redaction in AI pipelines, and network isolation of model and data endpoints.
  • Work with InfoSec and Group Architecture on threat modelling, security review, vulnerability and patch management, and remediation of findings against platform components.
  • Build compliance evidence into the platform: policy-as-code, drift detection, configuration baselines, and audit logging that stand up to internal and external review.
  • Write and maintain Terraform for the full estate: Foundry and AI services, Databricks, networking, identity and role assignments, Key Vault, storage, monitoring, and policy.
  • Own module structure, state management, workspace and environment separation, variable and secret handling, versioning, and drift detection for the infrastructure you build.
  • Build CI/CD pipelines in Azure DevOps or GitHub Actions covering plan and apply gates, automated testing, security and policy scanning, container builds, environment promotion, and rollback.
  • Publish reusable modules and reference implementations with documentation so other teams can consume the platform without asking you first.
  • Automate onboarding of new workloads and environments, so provisioning is a pipeline run, not a project.
  • Use coding agents and AI development tooling (Claude Code, GitHub Copilot, and similar) as a primary part of your daily workflow for implementation, refactoring, test generation, and debugging.
  • Build and maintain the scaffolding that makes agent-assisted development work on our repositories: context files, tool definitions, repository conventions, task decomposition patterns, and reusable prompt assets.
  • Apply engineering judgment to agent output. Review generated infrastructure and code as rigorously as human-written work and know where the tooling saves hours and where it creates cleanup work.
  • Automate repetitive platform works with scripted agents: module migrations, policy and documentation generation, dependency upgrades, and environment scaffolding.
  • Share working patterns with the team through examples in the codebase rather than through process documents.
  • Instrument the platform with Azure Monitor, Application Insights, and Log Analytics, including distributed tracing, token and cost telemetry, and structured logging across gateway, agent runtime, and data pipelines.
  • Build dashboards and alerts for latency, error rates, throughput, quota consumption, model usage, pipeline health, and spend.
  • Own FinOps for AI and data workloads: cost attribution by team and workload, budget alerts, rightsizing, reservation and commitment planning, and reporting that leadership can act on.
  • Define and meet availability, recovery, and backup expectations for platform services, including DR approach and tested restore procedures.
  • Participate in production support: triage, root cause analysis, fixes, and post-incident follow-through, and write runbooks so others can operate what you build.
  • Implement the technical controls behind Howden's AI SDLC and governance framework: environment gates, approval checkpoints, capability and agent registries, and evidence capture at deployment time.
  • Support architecture and security review by producing accurate as-built documentation, control mappings, and platform reference diagrams.
  • Contribute to platform standards and patterns alongside the Leads, Group Architecture, and InfoSec: bring options, trade-offs, cost and latency implications, and a recommendation grounded in what you have built.
  • Build spikes and reference implementations that prove or kill a design approach before the team commits to it.
  • Support product teams consuming the platform: onboarding, troubleshooting, pattern guidance, and unblocking integration work.
  • Work inside an Agile team with product managers, engineers, architecture, and security to plan, refine, and deliver iteratively.
  • Work across Howden US and Group functions, including Group Architecture, InfoSec, and infrastructure teams, to land shared platform decisions.
  • Review peer pull requests and raise the quality bar through code review rather than through process.
  • Flag design gaps, security exposure, and integration risks early to the Leads and enterprise architecture.
  • Write clear technical documentation for the platform and infrastructure you own.

Benefits

  • Medical, dental, and vision insurance, including healthcare savings and reimbursement accounts
  • 401(k) retirement plan
  • Flexible Paid Time Off and paid parental leave
  • Life and Disability insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service