Observability Analyst

State of OklahomaOklahoma City, OK
Onsite

About The Position

The Observability Analyst is responsible for day-to-day monitoring, triage, correlation, and first-line coordination with domain teams and the IT Operations Command Center (ITOCC). You will maintain and continuously improve the monitoring and observability capabilities that keep the organization's server, cloud, and network environments healthy. Using Datadog or similar platforms, this role builds meaningful dashboards and alerts and turns raw telemetry into early warning signals that let teams find and fix problems before they impact the business. You will focus on pattern analysis and hunting silent anomalies, feeding findings into the continuous-improvement loop. This is a hands-on technical role for someone who enjoys making complex environments visible, understandable, and measurably more reliable.

Requirements

  • Associate's or Bachelor's degree in Information Technology, Computer Science, or related field, or equivalent hands-on experience.
  • 2+ years of experience in infrastructure monitoring, observability, systems administration, or IT operations.
  • Hands-on experience with Datadog or a comparable observability platform (e.g., Dynatrace, New Relic, Splunk, Prometheus/Grafana).
  • Working knowledge of server, cloud (AWS, Azure, or GCP), and network fundamentals sufficient to instrument and troubleshoot across environments.
  • Experience configuring dashboards, alerts, and notification workflows that support fast, accurate incident response.

Nice To Haves

  • Familiarity with ITSM practices (incident, problem, change management) and on-call/escalation processes.
  • Exposure to AIOps or automated remediation approaches that reduce manual intervention.
  • Experience monitoring environments in regulated industries with attention to security and compliance requirements.

Responsibilities

  • Design, deploy, and maintain monitoring and observability tooling (Datadog or similar) across server, cloud, and network environments.
  • Instrument infrastructure, applications, and services with metrics, logs, traces, and synthetic checks to provide full-stack visibility.
  • Build and maintain dashboards that give clear, role-appropriate visibility into system health for engineers, managers, and leadership.
  • Configure and tune alerting thresholds and escalation policies to catch real issues early while minimizing noise and alert fatigue.
  • Integrate monitoring tools with incident management, ticketing, and on-call notification systems (e.g., PagerDuty, ServiceNow, Slack).
  • Track and report on key IT health indicators such as uptime, latency, error rates, capacity headroom, and patch/compliance status across multiple environments.
  • Partner with infrastructure, cloud, and network teams to identify recurring issues and drive root-cause fixes rather than repeated firefighting.
  • Support capacity planning by analyzing utilization trends and flagging environments approaching risk thresholds.
  • Contribute to post-incident reviews by providing telemetry, timelines, and health data that clarify what happened and why.
  • Continuously refine monitoring coverage as new systems, services, and cloud resources are added, retiring stale checks and dashboards.
  • Work closely with server, cloud, network, and application teams to understand what 'healthy' looks like for each environment and translate that into monitoring coverage.
  • Document monitoring standards, runbooks, and dashboard conventions so coverage stays consistent as the environment grows.
  • Train and support other engineers in interpreting dashboards, alerts, and observability data.
  • Evaluate and recommend improvements or additions to the observability toolset as monitoring needs evolve.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service