About The Position

The Node Connectivity team is responsible for a globally distributed service that connects millions of Meraki devices to the cloud services that manage them. We are looking for a technical lead to join our Kubernetes Site Reliability Engineering team. The initial focus of this role will be supporting the migration of the globally distributed service from AWS-hosted Kubernetes clusters to an internal, data-center-based Kubernetes platform. Because it has specialized networking requirements, including unique traffic patterns, load-balancing behavior, and cross-region connectivity, this work will help expand the SRE team’s ability to support specialized Kubernetes workloads at scale. Following the initial migration, the role will continue supporting the production operations while contributing to broader Kubernetes reliability, scalability, and platform engineering initiatives. The scope may evolve over time based on the needs of the Kubernetes SRE organization and the broader Cloud Network Platform Engineering organization.

Requirements

  • 10+ years of professional software, site reliability, or infrastructure engineering experience, including technical leadership of substantial production systems.
  • Experience designing, deploying, and operating large distributed services on Kubernetes.
  • 5+ years of programming experience in Go or a similar systems programming language.
  • Experience supporting production services through incident response, solving, performance analysis, Kubernetes reliability practices, observability, automation, and operational readiness.
  • Linux and networking concepts, including IPv4/IPv6, TCP, routing, DNS, and TLS; sound judgment in architecture, incident response, prioritization, and technical tradeoffs; and the ability to align stakeholders across teams without relying on formal authority.

Nice To Haves

  • Experience operating Kubernetes across AWS, on-premises, or hybrid environments, using AWS/EKS, Docker, Kustomize, GitLab CI/CD, or similar deployment systems.
  • Experience with Kubernetes networking, ingress, load balancing, service discovery, and traffic management for high-throughput or geographically distributed services.
  • Experience with gRPC, Protocol Buffers, mutual TLS, PKI, VPNs, tunneling, or network security.
  • Experience with production observability platforms such as OpenTelemetry, Prometheus, Datadog, or similar tools, along with distributed routing, packet processing, performance optimization, or failure testing.

Responsibilities

  • Lead the feasibility assessment, technical planning, and phased migration of eligible workloads from AWS-hosted Kubernetes clusters to the internal Kubernetes platform, establishing technical direction and staged delivery plans.
  • Support the operation and reliability of specialized Kubernetes workloads with non-standard load-balancing, networking, and traffic requirements.
  • Continue supporting the application after migration, including production readiness, incident response, performance and operational improvements.
  • Solve complex issues across applications, Kubernetes, Linux, networking, containers, and infrastructure; improve the scalability, reliability, security, performance, and operability of Kubernetes-hosted services.
  • Partner with the Node Connectivity, firmware, cloud infrastructure, security, SRE, and product teams to prioritize support and coordinate cross-system changes; contribute to broader Kubernetes SRE and platform reliability initiatives as priorities evolve.
  • Design, implement and maintain production-quality Go software for distributed, concurrent, and networked systems; Drive operational readiness through SLIs/SLOs, comprehensive testing, on-call support, root-cause analysis, and long-term corrective actions.

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday
  • paid year-end holiday shutdown
  • 4 paid days off for personal wellness determined by Cisco
  • 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees (non-exempt)
  • flexible vacation time off program (exempt)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter
  • up to 80 hours of unused sick time carried forward from one calendar year to the next
  • Additional paid time away may be requested to deal with critical or emergency issues for family members
  • Optional 10 paid days per full calendar year to volunteer
  • annual bonuses (for non-sales roles)
  • performance-based incentive pay (for sales roles)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service