Staff Infrastructure Engineer

The Hartford•Columbus, OH
•$116,800 - $175,200•Hybrid

About The Position

The Staff Infrastructure Engineer is responsible for leveraging agentic AI to quickly identify, assess, and resolve critical service issues and major incidents, ensuring swift detection and response. The successful candidate will demonstrate engineering-level technical depth in AI, cloud computing concepts, and infrastructure platforms, partnering with cross-functional teams for service restoration and continuously improving the event management process. This position is critical in enhancing enterprise resilience and implementing proactive AI-driven operations within the Technology Command Center. This role will have a Hybrid work schedule, with the expectation of working in an office (Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).

Requirements

  • 8+ years' experience in monitoring and observability tools such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes, or similar.
  • Experience in AI-driven data analysis and problem solving.
  • AI tool proficiency including prompt engineering and Interrogation skills
  • Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation, and documentation.
  • Understanding of event correlation, alert prioritization, service impact analysis, and basic automation/workflow enablement.
  • Strong analytical thinking and operational judgment.
  • Ability to work under pressure during high-impact events.
  • Candidate must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.

Nice To Haves

  • May require shift coverage, off-hours support, and escalation activities aligned to a 24x7 operational model and follow-the-sun support structure.

Responsibilities

  • Manage technical and executive communication for major incidents.
  • Perform alert triage, correlation, and initial impact assessment to separate actionable events from non-actionable noise.
  • Support major incident detection and escalation by validating symptoms, confirming affected services, and engaging resolver teams.
  • Use standard operating procedures, runbooks, and decision frameworks to investigate, prioritize, and escalate events.
  • Maintain situational awareness during active events; document event patterns and operational observations.
  • Promote events to incidents when thresholds are met; partner with internal and external teams during restoration activities.
  • Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness.
  • Collaborate with infrastructure, application, reliability engineering, service desk, and vendor teams to ensure effective event response and service restoration.
  • Support continuous improvement by identifying recurring alert issues, documenting operational insights, and recommending enhancements to monitoring, runbooks, and automation.

Benefits

  • short-term or annual bonuses
  • long-term incentives
  • on-the-spot recognition
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service