Infrastructure Engineer Lead – Cloud AI

American Electric Power•Gahanna, OH
•Onsite

About The Position

This hands-on role builds and operates secure, scalable infrastructure for AI and machine learning workloads across AWS and approved on-premises environments. Responsibilities include cloud foundations, networking, containers, compute, identity, cost management, and observability. The engineer partners with AI, architecture, cybersecurity, network, and application teams to deliver secure, cost-effective, production-ready infrastructure while supporting core cloud engineering services.

Requirements

  • Demonstrated hands-on engineering experience in AWS, including IAM, VPC networking, compute, storage, encryption/KMS, and account/organization structure.
  • Strong Kubernetes expertise including cluster design, operations, troubleshooting and security in EKS, ROSA or OpenShift.
  • Solid networking fundamentals, including routing, DNS, load balancing, traffic management, firewalls, hybrid connectivity, and network design for disconnected or intermittently connected environments.
  • Experience with API gateways, integration patterns, and exposing and securing services across environments.
  • Proficiency with infrastructure-as-code and automation (Terraform required; Ansible, Python or PowerShell preferred) and CI/CD pipelines.
  • Working knowledge of cloud security principles, identity and access management, and compliance requirements.
  • Experience implementing monitoring and cost management for cloud environments, with the ability to establish local monitoring, logging, patching, backup, and recovery processes for on-premises and disconnected environments.
  • Strong analytical, troubleshooting and problem-solving skills, with the ability to work independently on complex assignments.
  • Effective written and verbal communication skills, including the ability to present technical recommendations clearly to management and non-technical stakeholders.
  • Adhere to policies, procedures, standards, codes and regulations relevant to assignments.
  • Demonstrate in-depth knowledge of AEP infrastructure, environment and components to enable efficient, comprehensive responses to projects and problems.
  • Participate in on-call rotation, after-hours maintenance windows, and storm/emergency response support as required.

Nice To Haves

  • Architecture experience including designing end-to-end cloud solutions and influencing platform direction is highly desirable.
  • Direct experience building or operating AI/ML or GenAI workloads across cloud or on-premises infrastructure (Amazon Bedrock, SageMaker, vector databases, retrieval-augmented generation patterns, GPU-based training or inference, edge AI, or disconnected environments).
  • Experience with FinOps tooling and practices for high-variability workloads.
  • Experience in a regulated industry (utility, energy, financial services) or with FedRAMP/GovCloud environments.
  • Familiarity with multi-cloud environments (Azure, OCI, GCP) and enterprise data platforms such as Snowflake.
  • Certifications: AWS Solutions Architect – Associate or Professional; AWS Certified Machine Learning; Certified Kubernetes Administrator (CKA).

Responsibilities

  • Design, build and maintain the AWS account structure, landing zones, and reference patterns that host AI and machine learning workloads, leveraging AWS Control Tower, Landing Zone Accelerator, and Terraform.
  • Engineer the compute, storage, and networking foundations required for AI workloads, including GPU and accelerated instance families, high-throughput storage, and data access paths to enterprise data platforms.
  • Enable and operate AI platform services such as Amazon Bedrock and SageMaker, including private connectivity, model access provisioning, logging, and quota management.
  • Build reusable infrastructure-as-code modules, templates and pipelines so that AI teams can deploy quickly within approved guardrails.
  • Design, build and maintain on-premises AI infrastructure for edge and specialized use cases.
  • Support AI in disconnected or intermittently connected environments, including local compute, model distribution, patching, monitoring, backup, and recovery.
  • Create deployment standards, automation, runbooks, and support processes for on-premises and edge AI.
  • Integrate on-premises AI with enterprise identity, security, networking, monitoring, and governance where feasible.
  • Partner with Cybersecurity, Enterprise Architecture and Compliance to align AI infrastructure with AEP security standards, regulatory obligations, and responsible AI guardrails.
  • Review and remediate configuration drift, vulnerabilities and audit findings across the AI cloud estate; support evidence requests and control attestations.
  • Contribute to onboarding and intake processes for new AI use cases, ensuring workloads are provisioned into the right accounts with the right controls from day one.
  • Establish and operate FinOps practices for AI workloads, including tagging standards, showback/chargeback, budgets, anomaly detection, and forecasting.
  • Analyze and optimize spend on GPU compute, inference and token consumption, storage, and data transfer; recommend commitment strategies such as Savings Plans and Reserved Instances.
  • Provide cost transparency and consumption reporting to business stakeholders and technology leadership, and identify optimization opportunities before they become budget issues.
  • Partner with monitoring team to create logging, alerting and dashboards for AI infrastructure and workloads.
  • Define service-level objectives and operational thresholds for AI platforms, including model endpoint availability, latency, throughput, and error rates.
  • Support incident response, root cause analysis and problem management for AI-related infrastructure events; drive preventive actions and automation to reduce recurrence.
  • Perform standard cloud engineering functions across the AEP AWS environment including account provisioning, environment builds, platform upgrades, patching, automation and lifecycle management.
  • Design, deploy and operate Kubernetes environments (Amazon EKS, ROSA/OpenShift) including cluster architecture, autoscaling, node group and GPU scheduling, ingress, service mesh, RBAC, and cluster security hardening.
  • Engineer and support platform services including load balancing, traffic management, API gateways, and integration patterns across cloud, on-premises, edge, and disconnected environments.
  • Apply networking fundamentals, VPC design, subnetting, routing, Transit Gateway, Direct Connect, VPN, firewalls, TLS and certificate management to deliver secure, performant connectivity for AI and general workloads.
  • Prepare cost estimates, justifications, alternative solutions and technical recommendations; produce technical documentation, runbooks and standards.
  • Collaborate with Project Managers, Architects, Solution Engineers, Business Analysts and vendor partners to deliver consistent, reliable solutions that leverage AEP's technology standards, architectures and best practices.
  • Adhere to and advocate for change, incident and problem management processes; participate in on-call rotation and after-hours support as required.
  • Provide training, mentoring and technical work direction to other engineers on the team.

Benefits

  • Base salary
  • Annual bonus
  • Long-term incentive
  • 401(k) match
  • AEP Pension
  • Comprehensive benefits package designed to support and enhance the overall well-being of employees.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service