Site Reliability Engineer

LSEGSt. Louis, MO

About The Position

We are hiring a Site Reliability Engineer to help us run and improve LSEG's internal observability platform. Our platform brings telemetry, dashboards, alerts, service health, and operational insight into one common experience. We use it to help engineering and service teams detect issues earlier, reduce incident impact, and improve operational reliability. The focus of this role is the platform itself - its reliability, supportability, and usability - and the standards and self-service patterns that help teams adopt it consistently. This is not a general application monitoring role; ownership of dashboards, alerts, or telemetry for every application team is not what we are looking for here.

Requirements

  • Experience supporting production platforms or services in an SRE, platform engineering, infrastructure, DevOps, or operations engineering role.
  • Experience using observability data such as metrics, logs, traces, alerts, dashboards, or service health views to investigate issues.
  • Understanding of incident response, problem management, service readiness, or operational support processes.
  • Experience using automation or infrastructure-as-code practices to support repeatable delivery and operations.
  • Solid understanding of cloud, container, Linux, networking, or distributed system environments.
  • Ability to communicate technical information clearly to engineering, operations, and service stakeholders.
  • Experience writing or maintaining operational documentation such as runbooks, support guides, or onboarding material.
  • A practical approach to improving reliability, reducing manual work, and helping teams use shared platforms optimally.

Nice To Haves

  • Experience with observability, monitoring, telemetry, or data pipeline technologies such as OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, Flink, or similar tools.
  • Experience with GitOps workflows and tools such as Git, CI/CD pipelines, pull requests, environment promotion, or configuration-as-code.
  • Experience building or supporting internal platforms used by multiple engineering teams.
  • Experience defining or using SLOs, SLIs, error budgets, alert quality measures, or service health models.
  • Experience supporting telemetry pipelines, data routing, data filtering, retention, or cost management.
  • Experience working in a regulated, financial services, or large enterprise technology environment.
  • Experience helping engineering teams adopt shared standards, templates, or self-service platform capabilities.

Responsibilities

  • Monitor platform health, investigate and diagnose issues, and support service recovery across production and non-production environments.
  • Contribute to service readiness through dashboards, SLOs, SLIs, runbooks, resilience checks, and performance validation.
  • Use metrics, logs, traces, and service health data to understand issues and guide practical decisions.
  • Support incident response, triage, problem management, and root cause analysis.
  • Drive follow-up actions that reduce repeat issues and improve operational consistency.
  • Improve automation and GitOps-based workflows to reduce manual work across the team.
  • Maintain documentation, onboarding guides, and self-service tooling so teams can operate with less reliance on direct support.
  • Provide direct enablement on telemetry standards, alerting guidance, SLO practices, and platform workflows.

Benefits

  • healthcare
  • retirement planning
  • paid volunteering days
  • wellbeing initiatives
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service