Senior Site Reliability Engineer

FiservSunnyvale, CA
Onsite

About The Position

You will join our global team in Sunnyvale and help operate financial platforms at scale. You will partner with cross-functional teams to improve reliability, automate operations, and drive continuous improvement across our cloud-native environments.

Requirements

  • Solid practical experience in site reliability, operations or DevOps at a mid-to-senior level.
  • Strong shell scripting skills and a foundation in programming concepts.
  • Hands-on experience with cloud workloads—specifically Google Cloud Platform (GCP) and GKE.
  • Proven experience with containerisation and orchestration (Kubernetes).
  • Working knowledge of Infrastructure as Code and configuration management (Terraform, Ansible, Puppet).
  • Familiarity with monitoring and observability tooling such as Prometheus, Grafana and Datadog.
  • In-depth understanding of HTTP(s) traffic, routing and load-balancing, with practical experience observing and operating HAProxy.
  • Comfortable using GitHub and GitHub Actions for code management, automation and IaC pipelines.
  • Strong troubleshooting skills, a pragmatic problem-solving approach and effective communication for cross-team collaboration.

Nice To Haves

  • Experience programming in Python, Go or Java.
  • Exposure to large-scale financial services platforms or highly regulated environments.
  • Experience defining SLIs/SLOs and managing error budgets in production environments.

Responsibilities

  • Design, build and maintain automation to eliminate manual, repetitive operational tasks (runbooks, deployment pipelines, remediation scripts).
  • Operate and enhance monitoring, logging and alerting systems to ensure strong observability across services.
  • Participate in on-call rotations and lead incident response activities; run and document post-incident RCA and follow-up actions.
  • Collaborate with stakeholders to define SLIs and SLOs, manage error budgets and translate reliability goals into measurable actions.
  • Forecast capacity needs and contribute to resource planning to ensure performance and cost-efficiency.
  • Troubleshoot production issues: deep-dive analysis, isolate root causes and implement durable fixes.
  • Drive continuous improvement and platform hardening through runbook improvements, automation, and best-practice adoption.
  • Work closely as part of an international team to deliver project outcomes and operational excellence.

Benefits

  • Annual incentive opportunity which may be delivered as a mix of cash bonus and equity awards
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service