Staff Production Engineer (SRE) (Federal)

ZscalerShort Hills, NJ
$119,000 - $170,000Hybrid

About The Position

We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions. You will debug live production issues at the OS and network level, write production-grade automation, and drive platform resilience.

Requirements

  • US Citizenship is required (due to the nature of assigned customers)
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms.
  • Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation.
  • Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies).
  • Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump
  • Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain

Nice To Haves

  • Hands-on experience operating and managing FreeBSD / BSD operating systems in production.
  • Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic environments.
  • Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis.

Responsibilities

  • Partner with Engineering and Networking teams to maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks.
  • Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb).
  • Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Bash.
  • Operate end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise.
  • Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, and strict CI/CD validation prior to production rollouts.

Benefits

  • Various health plans
  • Time off plans for vacation and sick time
  • Parental leave options
  • Retirement options
  • Education reimbursement
  • In-office perks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service