AI-Enabled Manager (Monitoring & Workload)

V2X•UNAVAILABLE, UNAVAILABLE
•Onsite

About The Position

The AI-Enabled Manager, Monitoring & Workload Health is a technical leadership role responsible for the continuous observability, performance, and resilience of the enterprise IT ecosystem. Reporting to the Director of Adaptive Infrastructure Network & Security, this leader will drive the transition from reactive infrastructure monitoring to proactive, self-healing operations. By leveraging agentic AI alongside enterprise monitoring platforms such as Microsoft System Center Operations Manager (SCOM) and SolarWinds Orion, the Manager will optimize the health of cloud workloads, on-premises infrastructure, network devices, and SaaS applications. This role is critical to establishing a resilient AIOps environment where predictive failure analysis and automated remediations eliminate operational downtime and reduce manual toil for the engineering teams.

Requirements

  • Bachelor’s degree in Information Technology, Computer Science, Systems Engineering, or a related field; OR an equivalent combination of education and experience from which comparable knowledge and job skills can be obtained. (One year related experience may be substituted for one year of education, if degree is required).
  • 5-10 years of experience designing and operating complex enterprise monitoring, infrastructure, or reliability engineering solutions.
  • Proven experience as a manager or team leader overseeing technical personnel and large-scale automation projects.
  • U.S. Citizenship
  • Deep technical expertise in administering and scaling Microsoft SCOM and SolarWinds Orion in enterprise environments.
  • Strong proficiency in network automation, Infrastructure-as-Code (IaC), and scripting languages (PowerShell, Python, YAML, JSON) to drive automated remediation workflows.
  • Experience integrating AIOps tools and agentic AI models with ITSM platforms (e.g., ServiceNow) for closed-loop incident resolution.
  • Understanding of modern cloud architectures (Azure, AWS), virtualization (VMware, Hyper-V), and Zero Trust networking principles.
  • Exceptional collaboration skills to partner with service desk, security, and application teams to build a cohesive, automated operational fabric.

Responsibilities

  • Design, deploy, and manage enterprise monitoring solutions, specifically focusing on optimizing SolarWinds Orion and Microsoft SCOM across a hybrid IT environment.
  • Integrate network telemetry, system logs, and application performance data into a centralized AIOps platform to support intelligent operations and anomaly detection.
  • Define and enforce monitoring baselines, thresholds, and intelligent alerting rules that prioritize business-critical SaaS applications and AI workloads.
  • Develop observability strategies that provide end-to-end visibility into complex infrastructure paths, including Zero Trust network segments and multi-cloud environments.
  • Implement agentic AI workflows to identify degradation patterns and execute predictive failure analysis before system outages occur.
  • Build and maintain automated remediation runbooks utilizing tools like Ansible, Terraform, and Python to allow AI agents to execute self-healing actions on infrastructure and network devices.
  • Establish strict governance and human-in-the-loop oversight thresholds for agent-driven automated remediations, ensuring safe and compliant execution within the production environment.
  • Lead post-incident reviews (RCA) leveraging AI-generated timelines and telemetry data to continuously tune predictive algorithms and prevent recurring issues.
  • Monitor AI model inference traffic, LLM API calls, and agent-to-agent communication, ensuring the underlying infrastructure meets required Quality of Service (QoS) and latency SLAs.
  • Collaborate with infrastructure engineering and cloud teams to right-size compute workloads based on automated capacity planning and performance trend analysis.
  • Develop and deliver performance-based KPIs to IT leadership, highlighting system uptime, mean time to remediate (MTTR), and the effectiveness of automated resolution rates.
  • Manage the day-to-day tasking, performance, and effectiveness of a team of monitoring engineers and AIOps developers.
  • Foster a culture of automation-first thinking, guiding the team to codify operational policies and build "Compliance-as-Code" and "Self-healing" network capabilities.
  • Manage enterprise vendor and supplier agreements related to monitoring platforms, ensuring technical debt is minimized and platforms are optimized for value.

Benefits

  • Healthcare coverage
  • Retirement plan
  • Life insurance, AD&D, and disability benefits
  • Wellness programs
  • Paid time off, including holidays
  • Learning and Development resources
  • Employee assistance resources
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service