Site Reliability Engineer – Network Observability

Team Red DogRedmond, WA
Remote

About The Position

Team Red Dog is hiring a Site Reliability Engineer – Network Observability for our client, a leading cloud and software provider and intelligent cloud leader. This role will operate and maintain enterprise-scale network observability platforms supporting Azure infrastructure, with a strong focus on Linux administration, syslog-ng, SNMP telemetry, security compliance, patching, and platform reliability. You will automate operational workflows using Ansible, Bash, PowerShell, and Python while troubleshooting complex network and systems issues and supporting highly available observability services at global scale. This is an opportunity to work with large-scale network telemetry and help connect observability data with AI-driven workflows for performance, capacity management, and security.

Requirements

  • Syslog-ng – 3+ years of hands-on experience operating, configuring, troubleshooting, patching, and maintaining syslog-ng in enterprise Linux environments.
  • Linux Systems Administration – 5+ years administering Linux systems, including security and OS patching, configuration management, capacity planning, troubleshooting, and production support.
  • Network Engineering & Observability – 3+ years of enterprise networking experience with strong knowledge of network monitoring and telemetry protocols including SNMP, SNMP Traps, NetFlow, and gNMI.
  • Automation & Configuration Management – Working knowledge of Ansible and playbooks , with experience automating operational tasks using Bash, PowerShell, or Python.
  • Bachelor's degree in computer science, computer engineering, a related technical field, or equivalent professional experience.
  • 5–7 years of enterprise experience in IT systems, network engineering, site reliability engineering, or a closely related role.
  • 5+ years of Linux systems administration experience.
  • 3+ years of hands-on syslog-ng experience.
  • 3+ years of enterprise network engineering experience.
  • Strong understanding of network observability and telemetry technologies including SNMP, SNMP Traps, NetFlow, and gNMI.
  • Strong knowledge of enterprise networking and routing and switching protocols.
  • Hands-on experience with Microsoft Azure or a comparable cloud platform.
  • Experience with system capacity planning, functional configuration, auditing, and capacity analysis tools.
  • Working knowledge of Ansible and Ansible playbooks.
  • Experience automating tasks with Bash, PowerShell, and/or Python.
  • Proficiency with regular expressions.

Nice To Haves

  • Experience with IBM SevOne Network Performance Manager, Broadcom AppNeta, or similar enterprise observability platforms is preferred.
  • Experience with source control platforms and DevOps practices.
  • Intermediate knowledge of data retrieval and query languages such as KQL and T-SQL is preferred.
  • Exposure to commercially available AI platforms is preferred.
  • Deep, hands-on experience operating network monitoring and observability systems rather than solely working with general cloud infrastructure.
  • Strong Linux administration skills combined with practical syslog-ng and trapd expertise.
  • Solid understanding of enterprise routing, switching, and network management protocols.
  • Candidates who can pair this operational depth with Ansible playbooks and scripting automation.
  • Experience with IBM SevOne, Broadcom AppNeta, Azure, and large-scale enterprise network telemetry environments.

Responsibilities

  • Administer and operate Windows and Linux virtual machines hosted in Azure, maintaining compliance with security and configuration standards.
  • Operate and maintain network observability platforms, with a primary focus on syslog-ng and trapd running on Linux.
  • Plan and execute security, operating system, and application patching and upgrades.
  • Maintain observability systems through capacity planning, configuration management, functional audits, and regular maintenance.
  • Author and maintain rules using regular expressions to support telemetry processing and platform operations.
  • Support IBM SevOne Network Performance Manager and Broadcom AppNeta observability platforms.
  • Investigate automated monitoring alerts and customer-reported incidents involving network observability platforms.
  • Troubleshoot complex network observability configurations, applications, operating systems, and infrastructure issues.
  • Assist network and security engineers with identifying traffic patterns, resource utilization, and performance trends.
  • Automate operational and administrative tasks using Bash, PowerShell, Python, and configuration management tools.
  • Deploy, configure, and manage Azure cloud services with an emphasis on scalability, reliability, security, and cost effectiveness.
  • Apply DevOps practices including CI/CD pipelines, infrastructure as code, and source control to streamline deployment and operational processes.
  • Perform business continuity and disaster recovery failover testing as appropriate.
  • Manage assigned projects and program components to deliver services against established objectives and timelines.
  • Participate in the team's on-call DRI rotation and support timely incident resolution.

Benefits

  • Health insurance (medical, dental, vision, and life)
  • Employer-matched 401K plan
  • Generous Paid Time Off
  • Flexible Paid Holiday Benefit
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service