About The Position

We are seeking a Director of Infrastructure to lead data center strategy, cloud/hosting modernization, and provisioning automation as part of a major enterprise technology transformation. This is an on-site role located at one of the locations referenced above. This role sits at the intersection of physical infrastructure, cloud enablement, and IT risk/control automation owning the resiliency, scalability, and regulatory readiness of our hosting environment while driving the shift from legacy, manual infrastructure processes to a zero touch, automated model. A core part of this role is building the infrastructure foundation for AI adoption from GPU/accelerator capacity and data platform readiness to AI driven operations and governance and partnering with AI/ML, Data, and Risk teams to make sure infrastructure keeps pace with the organization's broader AI transformation. The ideal candidate has led large scale data center and hosting transformations in a regulated (financial services or similarly compliance heavy) environment, has direct exposure to standing up AI/ML infrastructure, and is comfortable operating within a formal program governance model, reporting to executive leadership and coordinating across multiple work streams.

Requirements

  • 10+ years of experience in infrastructure, data center, or platform engineering, including 4+ years in a leadership role.
  • Demonstrated experience leading data center strategy regional expansion, high availability/resiliency architecture, and migration planning.
  • Strong background in hosting and cloud strategy (public cloud plus on prem/data center), containerization, and infrastructure as code.
  • Experience modernizing provisioning/path to production processes, including automation of manual, high volume infrastructure requests.
  • Experience operating in a regulated industry (financial services, healthcare, or similar) with exposure to technology risk/control frameworks and regulatory exam processes.
  • Track record managing third party/vendor relationships, contracts, and due diligence for infrastructure or hosting services.
  • Experience operating within a formal transformation or PMO/governance structure, with executive level reporting.
  • Experience planning or scaling infrastructure to support AI/ML workloads (compute capacity, data platform readiness, or MLOps tooling), or leading infrastructure's role in an enterprise AI adoption effort.
  • Excellent communication skills, including presenting to executive leadership and external partners/vendors.
  • Bachelor's degree in computer science, Engineering, or related field (or equivalent experience).

Responsibilities

  • Own the multi region data center strategy, including capacity planning, geographic diversification, and expansion (new builds, leases, and buildouts) to support growth and regulatory resiliency requirements.
  • Design and implement high availability architecture with metro synchronous replication, geographic/asynchronous failover, and multi fault zone coverage to eliminate single points of failure.
  • Lead facility and platform migration planning (migration factories, facility disposition) as new data centers come online.
  • Manage long lead time (core infrastructure) and short lead time (hosting/compute) procurement planning in coordination with sourcing and vendor teams.
  • Define and execute hosting strategy across data center and cloud, including standing up or maturing a Cloud Center of Enablement (CCoE) and cloud governance framework.
  • Own infrastructure design patterns and low-level designs to ensure adaptable, scalable, and secure hosting solutions.
  • Drive container modernization and platform strategy (CaaS, service mesh) to improve portability and reduce legacy platform dependency.
  • Lead infrastructure repave/provisioning modernization migrating from legacy, manual provisioning to automated, near zero touch pipelines (including third party provisioning/storefront tooling).
  • Partner with software delivery, release engineering, and DevOps teams to integrate infrastructure automation into a single, traceable path to production toolchain.
  • Reduce manual request volume and capacity overallocation through automation of infrastructure deployment, configuration, and drift control.
  • Improve developer/platform team experience through self-service infrastructure and standardized automation onboarding.
  • Own the infrastructure roadmap for AI/ML workloads GPU/accelerator capacity planning, high-performance networking and storage, and data center power/cooling requirements to support AI at scale.
  • Partner with Data, AI/ML Engineering, and Platform teams to build and scale infrastructure for model training, fine tuning, and inference (on prem, cloud, and hybrid).
  • Evaluate and integrate AI driven infrastructure operations tooling predictive capacity planning, anomaly detection, self-healing automation/runbooks, and intelligent monitoring/drift control to reduce manual toil across the provisioning and path to production pipeline.
  • Support enterprise GenAI initiatives (e.g., internal knowledge search/assistants, automation copilots) by ensuring the underlying infrastructure, data pipelines, and access controls are secure, compliant, and scalable.
  • Work with Risk and Compliance to extend IT risk/control frameworks to cover AI specific risks (model governance, data lineage, access to sensitive training data, third party AI vendor risk).
  • Stay current on AI infrastructure trends and vendor landscape (compute, orchestration, MLOps tooling) and advise executive leadership on build vs. buy and platform strategy decisions.
  • Partner with Risk, Compliance, and Internal Audit to map infrastructure controls to regulatory/control frameworks, close identified gaps, and increase the ratio of automated to manual controls.
  • Support regulatory exam preparation and response (e.g., resilience, technology risk) with accurate infrastructure narratives and evidence.
  • Own third party/vendor risk management for infrastructure and hosting suppliers, including due diligence, contract structuring, and ongoing performance/SLA management.
  • Establish metrics, dashboards, and governance cadence (aligned to a program/portfolio "control tower" model) to report infrastructure transformation progress to executive leadership.
  • Hire, mentor, and manage infrastructure, platform, and provisioning engineering teams.
  • Operate within a formal transformation program structure engaging with program/portfolio management, financial/contract management, and organizational change management functions.
  • Manage budget, vendor contracts, and cost/capacity tradeoffs for infrastructure investments.
  • Represent infrastructure work streams in executive and vendor/partner forums, including presenting roadmap, risk, and outcomes.

Benefits

  • Competitive compensation
  • Comprehensive insurance options
  • Matching contributions through the 401(k) plan and the share purchase plan
  • Paid time off for vacation, holidays, and sick time
  • Paid parental leave
  • Learning opportunities and tuition assistance
  • Wellness and Well-being programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service