Senior Site Reliability Engineer, Production Engineer - ThousandEyes(Hybrid)

CiscoSan Francisco, CA
$167,700 - $245,200Hybrid

About The Position

As a Senior Site Reliability Engineer (SRE), you will lead the design and management of large-scale, highly available distributed systems, collaborating with application teams to ensure the performance, reliability, and security of our SaaS platform. You will deploy resilient AWS cloud-native services and leverage CNCF-standard tools—such as Kubernetes, Prometheus, and Service Mesh—to standardize operations across our multi-region, microservice-based architecture. A core component of your work involves developing automation for service operations, including deployment, chaos testing, and "everything-as-code" strategies, to provide robust guardrails for our rapidly growing infrastructure. You will directly influence the ThousandEyes platform by identifying and resolving operational obstacles, ensuring our systems remain scalable under substantial daily data volumes. This role is exceptionally exciting because it places you at the center of mission-critical engineering, where you will solve complex scaling challenges and shape the future of our global platform's reliability.

Requirements

  • Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience, or equivalent related work experience
  • Proficiency in software development with languages such as Python or Go
  • Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle
  • Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols
  • Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs

Nice To Haves

  • Familiarity with procedures for operating a large-scale, highly available enterprise platform
  • Excellent communication and documentation skills
  • Strong sense of ownership, drive, and attention to detail
  • Expert-level knowledge of Kubernetes and its ecosystem
  • In-depth knowledge of cloud providers, preferably AWS

Responsibilities

  • Lead the design and management of large-scale, highly available distributed systems.
  • Collaborate with application teams to ensure the performance, reliability, and security of our SaaS platform.
  • Deploy resilient AWS cloud-native services.
  • Leverage CNCF-standard tools—such as Kubernetes, Prometheus, and Service Mesh—to standardize operations across our multi-region, microservice-based architecture.
  • Develop automation for service operations, including deployment, chaos testing, and "everything-as-code" strategies.
  • Identify and resolve operational obstacles.
  • Ensure systems remain scalable under substantial daily data volumes.

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday, paid year-end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco
  • 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees (non-exempt)
  • flexible vacation time off program (exempt)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter, and up to 80 hours of unused sick time carried forward from one calendar year to the next
  • Optional 10 paid days per full calendar year to volunteer
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service