Cloud Engineer - AI Gateway - Global Industrial

Genuine Parts Company•Atlanta, GA

About The Position

Under limited supervision, the AI Gateway Cloud Engineer III designs, deploys, and operates secure, scalable AI gateway platforms across cloud and Kubernetes environments. This role enables standardized access to approved AI models, agents, and services by automating platform provisioning, configuration, policy deployment, and release processes. The engineer implements routing, resiliency, traffic, token, caching, identity, encryption, data protection, and responsible AI controls to improve reliability, performance, security, and cost efficiency. This role establishes monitoring and operational practices, leads troubleshooting and root cause analysis, and manages platform capacity, availability, upgrades, backups, and disaster recovery. The AI Gateway Cloud Engineer III partners with security, cloud, data, architecture, and application teams to define onboarding standards, reference architectures, service-level objectives, and runbooks; evaluates emerging capabilities; manages priorities and continuous improvement; and mentors engineers in consistent platform engineering and governance practices.

Requirements

  • Typically requires a bachelor's degree and five (5) to eight (8) years of related experience or an equivalent combination.
  • Advanced knowledge of AI gateway architecture and cloud services across Google Cloud Platform, Microsoft Azure, or Amazon Web Services.
  • Hands-on experience deploying and operating highly available platform services in cloud and Kubernetes environments.
  • Experience integrating generative AI models, agents, and services through standardized gateway interfaces and reusable patterns.
  • Proficiency with infrastructure as code, declarative configuration, automated testing, and deployment pipelines.
  • Knowledge of model routing, load balancing, failover, rate limiting, token controls, caching, and usage-based cost optimization.
  • Experience securing AI services through authentication, authorization, encryption, secrets management, private connectivity, and policy-based controls.
  • Knowledge of prompt and response guardrails, content safety, data protection, auditability, and responsible AI governance.
  • Proficiency with monitoring, logging, tracing, alerting, usage metering, and cost analysis for distributed AI workloads.
  • Strong troubleshooting and root cause analysis skills across platform, network, provider, policy, and application integration layers.
  • Knowledge of capacity planning, service-level objectives, resiliency testing, upgrades, backup validation, and disaster recovery.
  • Ability to define reference architectures, onboarding standards, operational runbooks, and scalable platform engineering practices.
  • Ability to communicate and collaborate effectively with security, cloud, data, architecture, application, and business teams.
  • Experience with Python and one or more scripting languages, such as Linux shell or PowerShell, for platform automation and integration.
  • Ability to work independently, manage competing priorities, evaluate emerging capabilities, and mentor engineers.

Responsibilities

  • Designs, deploys, and operates secure, scalable AI gateway platforms across cloud and Kubernetes environments.
  • Integrates approved AI models, agents, and services through standardized gateway interfaces and reusable patterns.
  • Automates platform provisioning, configuration, policy deployment, and release processes.
  • Implements model routing, failover, rate limits, token controls, and caching to improve reliability, performance, and cost efficiency.
  • Secures AI traffic using identity, access, encryption, credential protection, private connectivity, and policy-based controls.
  • Applies prompt and response guardrails, content-safety policies, data protection, and responsible AI controls.
  • Establishes monitoring and alerting for requests, tokens, latency, errors, model usage, costs, and policy enforcement.
  • Troubleshoots platform, network, provider, and policy issues and leads root cause analysis for production incidents.
  • Manages platform capacity, availability, upgrades, resiliency testing, backups, and disaster recovery.
  • Defines onboarding standards, reference architectures, service-level objectives, and operational runbooks.
  • Partners with security, cloud, data, architecture, and application teams to enable compliant AI adoption.
  • Evaluates emerging AI gateway capabilities and recommends adoption based on security, scalability, supportability, and cost.
  • Manages project priorities, deliverables, operational commitments, and continuous improvement initiatives.
  • Mentors engineers and promotes consistent platform engineering, security, and governance practices.
  • Performs other duties as assigned.

Benefits

  • options for healthcare coverage
  • 401(k)
  • tuition reimbursement
  • vacation
  • sick
  • holiday pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service