Principal, AI Operations & Infrastructure Automation

IDC Research Inc.Boston, MA
$139,000 - $240,000Remote

About The Position

IDC is modernizing how infrastructure is run — shifting from manual, ticket-driven operations to an automated, AI-informed model that keeps pace with the speed of the business. We're looking for a Principal of AI Operations & Infrastructure Automation to lead this transformation: designing and standing up the automation, tooling, and operating model that will define how IDC's infrastructure team works going forward. This is a hands-on, high-velocity leadership role for someone who has actually built AIOps and infrastructure-as-code capability before — not just overseen it. You'll personally write and review automation, get into the tooling directly, and lead by example for the team you manage. You'll partner closely with the CIO, Cyber & Infrastructure leadership, and AWS as a strategic cloud and co-development partner to re-architect how infrastructure operations get done. This role builds and leads the team responsible for automation — setting technical direction, reviewing their work, and growing their skills as the operating model evolves.

Requirements

  • 8+ years in infrastructure, platform engineering, or IT operations, including 3+ years leading automation or AIOps initiatives.
  • Demonstrated track record of redesigning an operations function around automation — not just introducing point tools but changing how a team works.
  • Deep hands-on experience with cloud infrastructure automation (AWS strongly preferred), infrastructure-as-code (Terraform, CloudFormation, or similar), and CI/CD pipelines.
  • Working knowledge of AIOps and agentic automation platforms, and a clear point of view on where AI genuinely improves operational outcomes versus adds noise.
  • Practical automation skill in Python or an equivalent language, including building integrations against REST APIs and operational tooling.
  • Hands-on experience with observability and monitoring platforms — Datadog, Splunk, Dynatrace, New Relic, Prometheus/Grafana, Elastic, or similar — including instrumenting services and building useful alerting.
  • Strong program leadership skills: able to define a roadmap, sequence dependencies, and hit aggressive milestones.
  • Excellent executive communication skills — comfortable distilling complex technical transformation into clear, concise updates for senior leadership.

Nice To Haves

  • Experience operating in a regulated or compliance-sensitive environment (SOC 2, ISO 27001, or similar).
  • Prior experience partnering directly with a hyperscaler (AWS, Azure, or GCP) on joint automation or AI initiatives.
  • Background in site reliability engineering (SRE) practices and observability tooling.
  • Experience with agentic AI frameworks, RAG pipelines, or building internal AI copilots for technical teams.
  • Familiarity with ITSM platforms and process — ServiceNow, Jira Service Management, or similar — and comfort automating against them.
  • Familiarity with FinOps practices and cloud cost automation.
  • Relevant certifications (cloud architect or engineer, Kubernetes, ITIL v4, Terraform).

Responsibilities

  • Own the end-to-end redesign of the infrastructure operating model, moving core workflows (provisioning, monitoring, incident response, patching, capacity management) from manual execution to automated, self-healing pipelines.
  • Define a phased transformation roadmap with clear milestones, automation-coverage targets, and go-live gates — and drive execution against it at pace.
  • Lead change management for the infrastructure team through the transition — defining new roles, skills, and ways of working as automation is adopted.
  • Build the AIOps layer; evaluate, select, and deploy AIOps platforms and agentic tooling — anomaly detection, predictive alerting, automated remediation — suited to IDC's AWS-centric environment. Implement event correlation and noise reduction so a hundred alerts resolve into one actionable incident with a probable root cause attached.
  • Stand up LLM (large language model)-assisted triage that classifies, enriches, and routes incidents, pulling relevant runbooks, past resolutions, and change history into the ticket automatically.
  • Develop predictive capacity and failure models for critical services, and act on them before thresholds are breached.
  • Automate end to end, build and champion an infrastructure-as-code (IaC) standard across the organization, converting legacy manual processes into version-controlled, repeatable automation.
  • Deliver self-healing automation and closed-loop remediation for the highest-volume recurring incidents — restart, rollback, scale, reroute, and verify without a human in the loop. Automate service request fulfillment: access provisioning, environment setup, onboarding and offboarding, and routine change execution.
  • Own CI/CD (continuous integration/continuous delivery-deployment) pipelines for operational tooling, including testing and safe rollback for automations that touch production.
  • Directly manage and develop a team of infrastructure and automation engineers — hiring, coaching, and setting technical direction — while staying hands-on in the tooling and code yourself.
  • Partner with AWS and other strategic vendors to pilot and scale automation and AI-driven operations capabilities, including Bedrock/AgentCore-based tooling where relevant.
  • Establish new operational KPIs — automation coverage, mean time to detect and resolve, deployment frequency, toil eliminated — and report progress in concise, executive-ready formats.
  • Set guardrails for AI in operations: human-in-the-loop thresholds, blast-radius limits, audit logging, evaluation of model output quality, and clear rollback paths.
  • Ensure automated operations meet IDC's security, compliance, and resiliency standards, working closely with Cyber & Compliance.

Benefits

  • 15 vacation days (prorated based on start date)
  • 12 company-paid holidays
  • 6 paid sick days (prorated based on start date; may vary by state)
  • Medical, dental, and vision coverage
  • 2 floating holidays (prorated based on start date)
  • 1 volunteer day
  • 401(k) company match (IDC matches 3% on the first 6% of employee contributions)
  • Company-paid short-term disability
  • Company-paid life insurance
  • Company-paid parental leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service