AIOps & Observability Lead

Bessemer Trust•Woodbridge Township, NJ
•$140,000 - $180,000•Hybrid

About The Position

We are building a new Operations function and need a hands-on, technical leader to set the direction for our AIOps and observability strategy. This is a player-coach role, weighted toward player: someone who works directly in the tools, not just the roadmap. Our new Splunk Observability implementation is a foundation of a broader automated operational intelligence ecosystem. We will lean on managed-service partners to act as first-line "eyes on glass" for the alerts and events it generates — but this role owns the strategy, architecture, and quality behind what those partners are watching. The goal is not just detection; it's ensuring issues can be discovered, triaged, and resolved through tooling and automation, reducing L3 engineer escalations wherever possible. The role could expand over time into desktop/endpoint observability, including a Nexthink implementation.

Requirements

  • Strong, hands-on experience with observability, monitoring, and event management platforms (e.g., Splunk, Splunk ITSI, BigPanda, PagerDuty).
  • Proven experience designing and implementing AIOps or event-intelligence capabilities — alert correlation, noise reduction, and automation.
  • Experience defining incident management processes, including major incident response and escalation design.
  • Comfortable operating as both strategic lead and hands-on technical contributor — able to shape direction and work directly in the tools.
  • Scripting or automation experience (e.g., Python, PowerShell, REST APIs) to support event enrichment and automated remediation.
  • Strong communication skills, with the ability to align infrastructure, application, and operations teams around a shared observability strategy.

Nice To Haves

  • Experience with AWS or other cloud-native observability and automation tooling.
  • Experience with endpoint/desktop observability platforms such as Nexthink.
  • ITIL/ITSM familiarity, particularly incident, problem, and change management.
  • Experience managing or directing managed-service/outsourced monitoring partners.
  • Experience in regulated, high-availability, or large-scale enterprise environments.

Responsibilities

  • Define and own the observability and event management strategy, stack, and processes, including how major incidents are detected, triaged, and resolved.
  • Lead the rollout of Splunk Observability across our application and infrastructure portfolio, driving onboarding, instrumentation, and coverage expansion.
  • Lead implementation of event management tooling — such as, BigPanda, Splunk ITSI, or PagerDuty — to correlate, prioritize, and route alerts from Splunk Observability.
  • Coach and enable the Automation team to become proficient in the observability and event intelligence platform, including alert logic, dashboard creation, event enrichment, correlation workflows, triggered automations, MCP-style integrations, and runbook automation.
  • Partner with the Infrastructure Architect to build a roadmap toward first-class, enterprise-grade observability.
  • Develop and execute an AIOps strategy that turns telemetry into HITL-automated, self-healing operations — evaluating and building on tooling in Splunk, AWS, or elsewhere as appropriate.
  • Stand up and mature an AIOps/SRE capability: automated remediation, intelligent alert correlation, and reduced manual triage.
  • Direct managed-service partners providing "eyes on glass" monitoring — defining what they watch, how they escalate, and holding them accountable to SLAs.
  • Design escalation paths so that issues resolvable via tooling are handled there first, protecting L3 engineering time for what truly needs it.
  • Establish standards for event classification, severity, ownership, suppression, and closure across the environment.
  • Assess and help scope the expansion into desktop/endpoint observability, including a Nexthink implementation.
  • Ensure dashboards, alerts, and event workflows are built around operational decisions, not just technical visibility.
  • Align observability, event management, incident, problem, change, knowledge, and escalation workflows with ITIL/ITSM practices and ServiceNow processes.

Benefits

  • Competitive base salary plus discretionary annual bonus for select positions
  • A 401(k) plan with a generous annual profit-sharing contribution
  • Personalized development and career opportunities, including tuition reimbursement support
  • Comprehensive medical, dental, and vision plans with zero contributions for employee coverage
  • Employee assistance (EAP) and wellness programs
  • Hybrid work environment: 60% in office, 40% remote for most positions
  • Paid time off and paid parental leave
  • Employer-paid life insurance and short- and long-term disability coverage
  • Legal services and financial wellness plans at no cost to employees
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service