Senior Site Reliability Engineer, Production Engineer - ThousandEyes

CiscoSan Francisco, CA
$165,000 - $241,400Hybrid

About The Position

We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.

Requirements

  • 5+ years of experience in a related role
  • Proficiency in software development with languages such as Python or Go
  • Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle
  • Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols
  • Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs

Nice To Haves

  • Familiarity with procedures for operating a large-scale, highly available enterprise platform
  • Excellent communication and documentation skills
  • Strong sense of ownership, drive, and attention to detail
  • Expert-level knowledge of Kubernetes and its ecosystem
  • In-depth knowledge of cloud providers, preferably AWS

Responsibilities

  • Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • grants of Cisco restricted stock units
  • 10 paid holidays per full calendar year
  • 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday
  • paid year-end holiday shutdown
  • 4 paid days off for personal wellness
  • 16 days of paid vacation time per full calendar year (non-exempt)
  • flexible vacation time off program (exempt)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter
  • up to 80 hours of unused sick time carried forward
  • Optional 10 paid days per full calendar year to volunteer
  • annual bonuses (for non-sales roles)
  • performance-based incentive pay (for sales roles)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service