Senior Network Observability Engineer

St. Jude Children's Research HospitalMemphis, TN
$94,640 - $169,520

About The Position

The Enterprise Network Team is building out proactive observability across a large, multi-site network. This role owns visibility into network health, performance, and availability, turning raw telemetry into actionable insight before problems hit patients, staff, or research systems.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Business or related field of study required.
  • Four (4) years of IT experience with experience in infrastructure operations and engineering environments.
  • Experience in infrastructure design, systems analysis, and security management.
  • Some experience working with commercial models around infrastructure capacity procurement.
  • Experience in customer service and vendor management.
  • Some experience in business stakeholder engagement and management.

Nice To Haves

  • Master's degree preferred.
  • Hands-on experience with at least one of StatSeeker, ThousandEyes, SolarWinds, LogicMonitor, or similar preferred.
  • Strong understanding of SNMP, NetFlow/IPFIX, streaming telemetry, and syslog architectures preferred.
  • Experience with Azure Event Grid or similar event routing/messaging services preferred.
  • Deep understanding of Cisco routing/switching, with comfort working across Catalyst, Nexus, and NDFC environments preferred.
  • Experience integrating network systems with ITSM platforms like ServiceNow preferred.
  • Familiarity with healthcare or other high-availability environments preferred.
  • CCNP or equivalent certification preferred.
  • Proven performance in earlier role/comparable role.

Responsibilities

  • Design and maintain observability pipelines across StatSeeker, ThousandEyes, Cisco Catalyst Center, and NDFC.
  • Build dashboards and alerting that surface SLA breaches, capacity trends, and anomalies across the network.
  • Correlate telemetry across the network stack (Catalyst, Nexus, F5, ISE, BlueCat) into a single operational picture.
  • Integrate monitoring data into ServiceNow for auto-ticketing and queue visibility.
  • Route and process telemetry events using Azure Event Grid to support real-time alerting and downstream integrations.
  • Develop synthetic monitoring and active testing for key network paths and services.
  • Lead root cause analysis on major incidents using historical telemetry and trend data.
  • Establish baselines and thresholds for network performance across sites and platforms.
  • Document observability standards and mentor the team on interpreting and acting on data.
  • Perform other duties as assigned to meet the goals and objectives of the department and institution.
  • Maintains regular and predictable attendance.

Benefits

  • St. Jude is an Equal Opportunity Employer
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service