Sr. Infra Ops – IT Engineer

The HartfordColumbus, OH
$116,800 - $175,200Hybrid

About The Position

The Sr. Infra Ops - IT Engineer is responsible for leveraging agentic AI to quickly identify, assess, and resolve critical service issues and major incidents, ensuring swift detection and response. The successful candidate will demonstrate engineering-level technical depth in AI, cloud computing concepts, and infrastructure platforms, partnering with cross-functional teams for service restoration and continuously improving the event management process. This position is critical in enhancing enterprise resilience and implementing proactive AI-driven operations within the Technology Command Center. This role will have a Hybrid work schedule, with the expectation of working in an office (Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).

Requirements

  • 8+ years' experience in monitoring and observability tools such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes, or similar.
  • Experience in AI-driven data analysis and problem solving.
  • AI tool proficiency including prompt engineering and Interrogation skills
  • Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation, and documentation.
  • Understanding of event correlation, alert prioritization, service impact analysis, and basic automation/workflow enablement.
  • Strong analytical thinking and operational judgment.
  • Ability to work under pressure during high-impact events.
  • Candidate must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.

Responsibilities

  • Manage technical and executive communication for major incidents.
  • Perform alert triage, correlation, and initial impact assessment to separate actionable events from non-actionable noise.
  • Support major incident detection and escalation by validating symptoms, confirming affected services, and engaging resolver teams.
  • Use standard operating procedures, runbooks, and decision frameworks to investigate, prioritize, and escalate events.
  • Maintain situational awareness during active events; document event patterns and operational observations.
  • Promote events to incidents when thresholds are met; partner with internal and external teams during restoration activities.
  • Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness.
  • Collaborate with infrastructure, application, reliability engineering, service desk, and vendor teams to ensure effective event response and service restoration.
  • Support continuous improvement by identifying recurring alert issues, documenting operational insights, and recommending enhancements to monitoring, runbooks, and automation.

Benefits

  • short-term or annual bonuses
  • long-term incentives
  • on-the-spot recognition
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service