Site Reliability Engineer Lead

Saxon GlobalPlano, TX

About The Position

This role requires a strong background in Site Reliability Engineering (SRE) or Production Engineering, with hands-on expertise in managing large-scale distributed messaging systems. The ideal candidate will have a deep understanding of system reliability, scalability, and high availability design, as well as messaging reliability patterns. Experience with observability tools, incident management, and scripting/automation is crucial. The role also involves managing Linux/Unix and Windows production environments and understanding event-driven architectures, messaging platform security, and vulnerability remediation.

Requirements

  • Strong experience in Site Reliability Engineering / Production Engineering
  • Hands-on expertise with IBM MQ (queue managers, clustering, channels, DLQ management)
  • Hands-on expertise with Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
  • Hands-on expertise with large-scale distributed messaging systems and runtime management
  • Deep understanding of system reliability, scalability, and high availability design
  • Deep understanding of messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
  • Deep understanding of incident management, root cause analysis, and problem management
  • Experience with observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
  • Experience with event and anomaly detection in high-volume systems
  • Strong scripting/automation skills (Shell, Python, PowerShell)
  • Experience managing Linux/Unix and Windows production environments
  • Knowledge of event-driven architecture and messaging-based integration patterns
  • Understanding of messaging platform security (TLS, certificates, channel auth, encryption)
  • Understanding of vulnerability remediation and risk mitigation in production systems
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service