AI Observability Engineer

NCR VoyixAtlanta, GA

About The Position

Observability Automation Engineer Role Summary We are seeking an Observability Engineer to design, implement, and support enterprise observability platforms while driving automation, AI, and Agentic AI initiatives. The ideal candidate will leverage intelligent automation, AI-powered analytics, and autonomous agents to improve service reliability, accelerate troubleshooting, and reduce operational overhead. Key Responsibilities Design, deploy, and maintain enterprise observability platforms for monitoring, logging, tracing, and alerting. Develop dashboards, KPIs, and service health metrics to provide actionable operational insights. Implement and optimize observability solutions using tools such as Splunk, AppDynamics, or Splunk Observability Cloud platforms. Automate operational processes, alert management, health checks, and incident response workflows using scripting and orchestration tools. Collaborate with engineering and operations teams to improve application performance, reliability, and scalability. Analyze incidents, identify root causes, and implement preventive measures through proactive monitoring and automation. Drive adoption of AI-powered observability capabilities, including anomaly detection, predictive analytics, event correlation, and intelligent alerting. Leverage Microsoft Copilot, Generative AI, and automation technologies to enhance troubleshooting, operational efficiency, and engineering productivity. Develop AI-assisted runbooks, knowledge bases, and self-healing solutions to reduce manual intervention and Mean Time to Resolution (MTTR). Participate in on-call support and major incident management activities as needed. Design and implement AI-driven observability solutions using telemetry, monitoring, logging, and distributed tracing platforms. Develop automated remediation, self-healing workflows, and operational runbooks using scripting, orchestration, and infrastructure-as-code tools. Build and integrate Agentic AI solutions that can autonomously analyze alerts, retrieve operational context, recommend actions, and execute approved remediation workflows. Leverage Microsoft Copilot and Generative AI tools to improve incident investigation, root-cause analysis, knowledge management, and engineering productivity. Implement AI-powered anomaly detection, event correlation, capacity forecasting, and predictive monitoring capabilities. Develop integrations between observability platforms and AI agents to automate repetitive operational tasks and improve MTTR. Collaborate with application, SRE, cloud, and platform teams to identify opportunities for AI-assisted operations and process automation.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
  • 3+ years of experience in Observability, Site Reliability Engineering (SRE), Infrastructure Engineering, or Operations Engineering.
  • Strong experience with monitoring, logging, tracing, and performance management platforms.
  • Proficiency in scripting and automation using Python, PowerShell, Bash, or similar languages.
  • Experience working with cloud platforms such as Microsoft Azure, AWS, or Google Cloud.
  • Knowledge of CI/CD pipelines, Infrastructure as Code (Terraform, Bicep, ARM, etc.), and DevOps practices.
  • Experience with Python, PowerShell, or similar languages for automation and API integrations.
  • Experience developing automation workflows and integrating observability platforms with cloud services and operational tools.
  • Familiarity with AI-assisted engineering practices and Copilot-enabled development workflows.
  • Strong analytical and troubleshooting skills.

Nice To Haves

  • Experience with Microsoft Copilot, Azure OpenAI, Copilot Studio, or other AI-powered engineering tools.
  • Knowledge of AIOps platforms and machine learning concepts related to observability.
  • Experience implementing OpenTelemetry standards and distributed tracing solutions.
  • Familiarity with Kubernetes, containers, and cloud-native monitoring architectures.
  • Experience building automated remediation and self-healing workflows.
  • Hands-on experience with Microsoft Copilot, Azure AI Services, Azure OpenAI, Copilot Studio, LangChain, Semantic Kernel, Agentic AI frameworks, or similar technologies.
  • Experience building AI agents, retrieval-based knowledge systems, AI-powered chatbots, or autonomous operational workflows.
  • Knowledge of RAG architectures, vector databases, prompt engineering, and AI governance best practices.
  • Familiarity with AIOps platforms and event intelligence solutions.

Responsibilities

  • Design, deploy, and maintain enterprise observability platforms for monitoring, logging, tracing, and alerting.
  • Develop dashboards, KPIs, and service health metrics to provide actionable operational insights.
  • Implement and optimize observability solutions using tools such as Splunk, AppDynamics, or Splunk Observability Cloud platforms.
  • Automate operational processes, alert management, health checks, and incident response workflows using scripting and orchestration tools.
  • Collaborate with engineering and operations teams to improve application performance, reliability, and scalability.
  • Analyze incidents, identify root causes, and implement preventive measures through proactive monitoring and automation.
  • Drive adoption of AI-powered observability capabilities, including anomaly detection, predictive analytics, event correlation, and intelligent alerting.
  • Leverage Microsoft Copilot, Generative AI, and automation technologies to enhance troubleshooting, operational efficiency, and engineering productivity.
  • Develop AI-assisted runbooks, knowledge bases, and self-healing solutions to reduce manual intervention and Mean Time to Resolution (MTTR).
  • Participate in on-call support and major incident management activities as needed.
  • Design and implement AI-driven observability solutions using telemetry, monitoring, logging, and distributed tracing platforms.
  • Develop automated remediation, self-healing workflows, and operational runbooks using scripting, orchestration, and infrastructure-as-code tools.
  • Build and integrate Agentic AI solutions that can autonomously analyze alerts, retrieve operational context, recommend actions, and execute approved remediation workflows.
  • Leverage Microsoft Copilot and Generative AI tools to improve incident investigation, root-cause analysis, knowledge management, and engineering productivity.
  • Implement AI-powered anomaly detection, event correlation, capacity forecasting, and predictive monitoring capabilities.
  • Develop integrations between observability platforms and AI agents to automate repetitive operational tasks and improve MTTR.
  • Collaborate with application, SRE, cloud, and platform teams to identify opportunities for AI-assisted operations and process automation.

Benefits

  • Offers of employment are conditional upon passage of screening criteria applicable to the job
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service