Senior AI DevOps Engineer (AI Ops / Platform Engineering) (Hybrid)

NTT DATA ServicesAtlanta, GA
$87,952 - $162,875Hybrid

About The Position

We are currently seeking a Senior AI DevOps Engineer (AI Ops / Platform Engineering) (Hybrid) to join our team in Atlanta, Georgia (US-GA), United States (US). We're looking for an experienced Senior AI DevOps Engineer to help build the next generation of AI-powered software delivery and cloud operations. In this role, you'll combine modern DevOps practices with Generative AI, LLM agents, Model Context Protocol (MCP), and intelligent automation to transform how engineering teams build, deploy, and operate software. You'll partner with Platform Engineering, DevOps, Security, SRE, and AI teams to design secure, scalable, cloud-native solutions that accelerate software delivery while improving reliability, observability, and operational efficiency. This is an opportunity to work on cutting-edge AI technologies that are redefining modern software engineering.

Requirements

  • 7+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, or Infrastructure Automation.
  • 4+ years designing and supporting enterprise CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD, or similar tools.
  • 3+ years of cloud engineering experience in AWS, Azure, or Google Cloud (AWS preferred).
  • Strong experience deploying and managing Kubernetes and Docker in production environments (EKS, AKS, or GKE).
  • Hands-on experience with Infrastructure as Code using Terraform, OpenTofu, Pulumi, Terragrunt, CloudFormation, or similar tools.
  • Strong programming skills in Python, TypeScript, JavaScript, Bash, Go, or similar languages.
  • Experience integrating enterprise LLM platforms such as OpenAI, Anthropic, or equivalent AI services into engineering workflows.
  • Experience with AI orchestration frameworks such as LangChain, CrewAI, LlamaIndex, or similar technologies.
  • Experience designing or implementing Model Context Protocol (MCP) clients and servers.
  • Experience implementing DevSecOps practices including SAST, DAST, dependency scanning, container security, secrets management, and vulnerability management.
  • Experience with secrets management solutions such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault.
  • Strong experience with monitoring, logging, and observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, Splunk, Dynatrace, or ELK.
  • Excellent troubleshooting skills across cloud infrastructure, CI/CD pipelines, Kubernetes, and production systems.
  • Strong communication skills and the ability to collaborate across engineering, security, and AI teams.

Nice To Haves

  • Experience building AI-assisted infrastructure provisioning and deployment workflows.
  • Experience implementing autonomous or AI-assisted incident response and operational remediation.
  • Experience with MLOps platforms including MLflow, Amazon SageMaker, Vertex AI, Azure ML, or similar technologies.
  • Experience implementing human-in-the-loop approval workflows for AI-generated operational actions.
  • Knowledge of Policy-as-Code frameworks such as Open Policy Agent (OPA), Sentinel, or Checkov.
  • Experience with GitOps platforms such as ArgoCD or Flux.
  • Experience working within regulated industries such as financial services, healthcare, insurance, or government.
  • AWS, Kubernetes, DevOps, Security, or AI/ML certifications.

Responsibilities

  • Design, build, and optimize AI-enabled CI/CD pipelines that improve developer productivity, deployment speed, and software quality.
  • Develop and deploy Model Context Protocol (MCP) clients and servers that securely connect enterprise LLMs with engineering tools, cloud infrastructure, and operational platforms.
  • Build custom MCP services using Python, TypeScript, JavaScript, or Node.js to expose infrastructure, deployment, monitoring, and operational data to authorized AI agents.
  • Integrate LLM-powered workflows for automated code reviews, testing, security analysis, release validation, and infrastructure recommendations.
  • Build and maintain enterprise CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD, or similar platforms.
  • Implement AI-driven ChatOps capabilities that enable engineers to interact with deployment pipelines, cloud environments, and operational tools through secure conversational interfaces.
  • Design intelligent remediation workflows for incident detection, root cause analysis, log analysis, and operational troubleshooting.
  • Develop secure Infrastructure-as-Code automation using Terraform, OpenTofu, Pulumi, Terragrunt, CloudFormation, or similar technologies.
  • Deploy and manage containerized applications using Kubernetes and Docker across AWS, Azure, or Google Cloud.
  • Build AI-powered observability solutions leveraging Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, ELK, or similar platforms.
  • Implement security guardrails including RBAC, least-privilege access, approval workflows, audit logging, rollback mechanisms, and secure AI tool access.
  • Partner with Engineering, Platform, Security, SRE, and AI teams to identify and implement intelligent automation opportunities.
  • Create reusable automation frameworks, documentation, dashboards, and engineering best practices that scale across the organization.

Benefits

  • medical, dental, and vision insurance with an employer contribution
  • flexible spending or health savings account
  • life and AD&D insurance
  • short- and long-term disability coverage
  • paid time off
  • employee assistance
  • participation in a 401k program with company match
  • additional voluntary or legally-required benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service