Observability Engineer

Indotronix International CorporationIrving, TX
Onsite

About The Position

Join a dynamic technology team as a Senior Observability Engineer based in Irving, TX. You will architect and lead end-to-end observability solutions across complex, hybrid environments, transforming platform telemetry into actionable insights that drive reliability and performance. This is an opportunity to shape observability strategy, work with cutting-edge tools, and collaborate with engineering, operations, and leadership. Advance your career by building and standardizing monitoring frameworks at scale.

Requirements

  • 7+ years in IT operations, SRE, systems engineering, or infrastructure
  • 5+ years designing and implementing observability/monitoring solutions across distributed systems
  • 5+ years production experience with Kubernetes platforms
  • Expertise in logging, tracing, and enhanced monitoring
  • Strong hands-on experience with Grafana (dashboards, alerts, data sources)
  • Advanced Splunk SPL, dashboarding, and investigation capabilities
  • Proficiency with querying languages (SQL, PromQL)
  • Deep understanding of Kubernetes internals, OpenShift (OCP), AKS, GKE
  • Experience with Prometheus, OpenTelemetry, Kubernetes exporters
  • OS-level monitoring (Linux/Windows) and network fundamentals
  • Experience with ThousandEyes, BigPanda, and ServiceNow ITSM workflows

Nice To Haves

  • Experience designing SLOs/SLIs, reliability scorecards
  • Familiarity with Istio, service mesh metrics, mTLS
  • Capacity planning and trend analysis using observability data
  • Exposure to multi-cloud observability strategies
  • Monitoring for databases, message brokers, middleware
  • Familiarity with AIOps or ML-driven anomaly detection

Responsibilities

  • Architect and implement comprehensive observability frameworks across cloud, on-premises, networking, databases, middleware, and applications
  • Evaluate, select, and integrate observability and monitoring tools, establishing reference architectures
  • Design and maintain standardized Grafana dashboards for platform and workload health (OCP, AKS, GKE)
  • Define golden signals and platform health KPIs tied to availability, performance, and reliability
  • Serve as an advanced Splunk user: develop complex SPL queries, dashboards, and root-cause investigations
  • Correlate logs, metrics, and events across Grafana and Splunk to drive rapid incident resolution (MTTR reduction)
  • Implement and tune platform-specific observability for Kubernetes platforms (OCP, AKS, GKE)
  • Configure and manage ThousandEyes for synthetic monitoring and network intelligence
  • Administer BigPanda for AIOps-driven event correlation and noise reduction
  • Integrate ServiceNow for automated incident creation and enriched alerting
  • Document observability standards, dashboards, and onboarding processes

Benefits

  • Work in a collaborative, technology-driven environment
  • Direct impact on reliability, performance, and operational excellence
  • Exposure to the latest observability and cloud-native technologies
  • Career growth opportunities in a high-visibility engineering role
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service