Senior Observability Engineer (SRO)

NextEra EnergyPlantation, FL

About The Position

The Senior Observability Engineer / Architect will help advance modern SRE practices across enterprise IT operations, with a focus on service-level visibility, observability, actionable alerting, automation, runbook maturity, and proactive service-health management. This role will partner across Observability, Event Management, infrastructure, application, and ServiceNow teams to improve reliability engineering standards, strengthen monitoring coverage, reduce operational noise, and support the transition from reactive incident response to data-driven, service-health operations. The position will provide a technical connection point across Information Technology, infrastructure, application, cloud, ServiceNow, and operations teams to improve reliability, observability, and operational readiness. This role will support the development and execution of observability standards, service health practices, alerting improvements, automation opportunities, and reliability engineering patterns that help teams detect, understand, and resolve service issues more effectively.

Requirements

  • Strong understanding of SRE principles and reliability engineering practices.
  • Experience with metrics, logs, traces, events, and dashboards.
  • Knowledge of observability tools such as ScienceLogic, ServiceNow ITOM, Splunk, or AppDynamics.
  • Ability to define SLIs, SLOs, and service-health indicators.
  • Experience improving alert quality, routing, and correlation.
  • Strong troubleshooting and problem-solving skills.
  • Scripting or automation experience to reduce manual effort.
  • Ability to collaborate across infrastructure, application, cloud, and operations teams.
  • High School Grad / GED
  • Bachelor's or Equivalent Experience
  • Experience: 4+ years

Nice To Haves

  • Bachelor's Degree

Responsibilities

  • Assist in designing, implementing, and operating enterprise observability capabilities across metrics, logs, traces, events, synthetic monitoring, dashboards, and service-health views.
  • Apply modern SRE principles, including SLIs, SLOs, error-budget thinking, toil reduction, automation, incident learning, and reliability-focused engineering practices.
  • Partner with infrastructure, application, cloud, database, network, storage, and operations teams to define monitoring requirements, alert thresholds, escalation paths, and service-health indicators.
  • Support observability platform capabilities across tools such as ScienceLogic, ServiceNow ITOM/Event Management, Splunk, cloud-native monitoring platforms, AppDynamics, synthetic monitoring, and related technologies.
  • Improve alert quality by helping ensure alerts are actionable, properly routed, associated with the correct configuration item or service, and supported by clear response guidance.
  • Assist in aligning operational events to ServiceNow Event Management, including event ingestion, alert correlation, suppression logic, incident creation criteria, and notification workflows.
  • Contribute to service-level visibility by supporting dashboards, scorecards, service maps, dependency views, ownership models, and operational health reporting.
  • Analyze recurring incidents, monitoring gaps, alert patterns, and operational trends to identify reliability improvement opportunities.
  • Develop and maintain runbooks, knowledge articles, technical documentation, monitoring standards, and operational handoff materials.
  • Support automation opportunities that reduce manual effort, improve triage consistency, and accelerate restoration while maintaining appropriate governance and controls.
  • Collaborate with teams during incidents and problem reviews to improve detection, escalation, root-cause analysis, and long-term prevention.
  • Help advance observability maturity through practical adoption of standards such as OpenTelemetry where appropriate, along with consistent telemetry collection and platform integration practices.
  • Respond to complex operational scenarios where standard procedures have not resolved the issue and provide technical analysis to support restoration and prevention.
  • Continuously evaluate observability practices, platform effectiveness, data quality, and service readiness to improve reliability outcomes across the enterprise.

Benefits

  • Wide range of benefits to support our employees and their eligible family members.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service