Observability Automation Engineer Role Summary We are seeking an Observability Engineer to design, implement, and support enterprise observability platforms while driving automation, AI, and Agentic AI initiatives. The ideal candidate will leverage intelligent automation, AI-powered analytics, and autonomous agents to improve service reliability, accelerate troubleshooting, and reduce operational overhead. Key Responsibilities Design, deploy, and maintain enterprise observability platforms for monitoring, logging, tracing, and alerting. Develop dashboards, KPIs, and service health metrics to provide actionable operational insights. Implement and optimize observability solutions using tools such as Splunk, AppDynamics, or Splunk Observability Cloud platforms. Automate operational processes, alert management, health checks, and incident response workflows using scripting and orchestration tools. Collaborate with engineering and operations teams to improve application performance, reliability, and scalability. Analyze incidents, identify root causes, and implement preventive measures through proactive monitoring and automation. Drive adoption of AI-powered observability capabilities, including anomaly detection, predictive analytics, event correlation, and intelligent alerting. Leverage Microsoft Copilot, Generative AI, and automation technologies to enhance troubleshooting, operational efficiency, and engineering productivity. Develop AI-assisted runbooks, knowledge bases, and self-healing solutions to reduce manual intervention and Mean Time to Resolution (MTTR). Participate in on-call support and major incident management activities as needed. Design and implement AI-driven observability solutions using telemetry, monitoring, logging, and distributed tracing platforms. Develop automated remediation, self-healing workflows, and operational runbooks using scripting, orchestration, and infrastructure-as-code tools. Build and integrate Agentic AI solutions that can autonomously analyze alerts, retrieve operational context, recommend actions, and execute approved remediation workflows. Leverage Microsoft Copilot and Generative AI tools to improve incident investigation, root-cause analysis, knowledge management, and engineering productivity. Implement AI-powered anomaly detection, event correlation, capacity forecasting, and predictive monitoring capabilities. Develop integrations between observability platforms and AI agents to automate repetitive operational tasks and improve MTTR. Collaborate with application, SRE, cloud, and platform teams to identify opportunities for AI-assisted operations and process automation.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level