Senior Site Reliability Engineer

2KBurnaby, BC
CA$93,700 - CA$138,700Hybrid

About The Position

The 2K SRE team owns the infrastructure behind every player connection — All 2K game services, account platforms, CI/CD pipelines, and developer tooling spanning AWS, GCP, and on-premises data centers across multiple global regions. Global launch windows and live-service events push systems to their limits, and this team is expected to hold the line. Post-mortems here focus on systems, not people. Automation is the default answer to repetitive work. The infrastructure keeps millions of players connected — and the team takes that seriously! The Senior SRE at 2K is a hands-on technical leader — shaping production infrastructure across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing the gap between what engineering ships and what players experience.

Requirements

  • 5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale
  • Deep Kubernetes experience in cloud environments (EKS or GKE preferred) — networking, storage, multi-cluster patterns
  • Strong IaC proficiency with Terraform and/or Pulumi; hands-on with Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions)
  • Modern and Legacy Tech: AWS, GCP, VMware, and Bare metal servers
  • Server Configuration using Ansible, Puppet, and AWS Systems Manager
  • Observability stack experience: Datadog, Prometheus + Grafana, and OpenTelemetry,
  • SLI/SLO/error budget fluency — including how to operationalize them inside engineering teams
  • Production-quality code in Go, Python, or TypeScript: tools, automation, and internal libraries
  • Linux internals, TCP/IP networking, DNS, and TLS — proven enough to debug at the system level
  • Incident response and post-mortem leadership with a track record of systemic follow-through

Nice To Haves

  • Live-service game or large-scale consumer internet experience at millions of concurrent users
  • Service mesh depth (Istio, Cilium) and advanced Kubernetes networking
  • FinOps and managing resources at cloud scale
  • Experience with AI and Agentic Development
  • Cloud certifications (AWS Solutions Architect, GCP Professional Cloud Architect, CKA/CKS, or equivalent)
  • Experience mentoring SREs or leading reliability working groups

Responsibilities

  • Design, build, and operate scalable multi-cloud and hybrid infrastructure using Terraform, Pulumi, and GitOps workflows (ArgoCD, Flux).
  • Own Kubernetes platforms (EKS, GKE) end-to-end — cluster lifecycle, multi-tenancy, networking (Istio, Cilium), and autoscaling — and push progressive delivery patterns (blue/green, canary) across game service deployments.
  • Build and run the full observability stack: Prometheus + Grafana + Datadog
  • Define SLI/SLO/error budget policies and build alerting that cuts through the noise
  • Lead chaos engineering exercises to surface failure modes before players encounter them
  • Drive incident response and post-mortems with a focus on systemic fixes and real follow-through
  • Eliminate toil through self-service provisioning, automated remediation, and intelligent scaling.
  • Harden CI/CD pipelines (GitHub Actions, Jenkins, ArgoCD).
  • Embed security at the platform layer through secrets management (PasswordState, 1Password, and AWS Secrets Manager), policy-as-code (OPA/Gatekeeper).
  • Promote SRE practices across 2K studios through reliability reviews, runbooks, and embedded collaboration
  • Shape architectural decisions and author engineering RFCs that move the platform forward

Benefits

  • bonus and/or equity awards
  • full range of medical, financial, and/or other benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service