Software Engineer II Platform Data Reliability

PlayStation GlobalSan Mateo, CA
$150,100 - $225,100Hybrid

About The Position

Ready to level up your career? Join PlayStation as a Software Engineer II focused on Platform Data Reliability and Automation and help build reliable, scalable experiences for millions of players around the globe. At PlayStation, we’re known not only for delivering exceptional gaming experiences but also for fostering an engineering environment centered on innovation, creativity, collaboration, and technical excellence. We welcome passionate engineers who enjoy solving challenging problems and are excited about shaping the future of play. We are seeking a Software Engineer II (Platform Data Reliability & Automation) to help build, automate, and operate scalable data platforms using Infrastructure as Code (IaC) and cloud technologies. This role focuses on improving the reliability and automation of NoSQL, streaming, and caching services across AWS and GCP environments. You’ll develop automation, observability, and operational tooling supporting technologies such as Cassandra, Aerospike, Kafka, and Redis. Working alongside senior engineers, platform teams, and product teams, you’ll contribute to highly available infrastructure supporting billions of transactions and millions of players globally. By applying software engineering and database reliability engineering principles, you’ll help reduce manual work, improve system uptime, and make data services easier and safer for engineering teams to use.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of experience in software engineering, database reliability engineering, site reliability engineering, platform engineering, or a related field.
  • Experience developing production software in Go, with an understanding of idiomatic code, testing, concurrency, error handling, and maintainability.
  • Hands-on experience developing or maintaining Infrastructure as Code or configuration-management tools such as Terraform or Ansible.
  • Experience deploying or operating workloads on Kubernetes.
  • Experience with AWS or GCP and familiarity with managed services such as MSK, DynamoDB, ElastiCache, Memorystore, or equivalent technologies.
  • Working knowledge of one or more NoSQL, caching, or streaming technologies, such as Cassandra, Aerospike, Kafka, AWS MSK, or Redis.
  • Understanding of distributed systems concepts, including availability, consistency, replication, fault tolerance, and horizontal scaling.
  • Working knowledge of Linux, networking, storage, and common system-troubleshooting techniques.
  • Familiarity with observability tools and practices, including metrics, logging, tracing, alerting, and dashboard creation.
  • Ability to independently diagnose and resolve technical problems within a defined scope, learn unfamiliar systems, and seek guidance when addressing complex or ambiguous challenges.
  • Strong written and verbal communication skills, with the ability to collaborate effectively across teams.

Responsibilities

  • Develop, maintain, and improve Infrastructure as Code and configuration-management automation using tools such as Terraform and Ansible to provision, configure, monitor, scale, and manage NoSQL, streaming, and caching platforms.
  • Build automation that enables repeatable and reliable deployment of data services across cloud and hybrid environments.
  • Contribute to the reliability, availability, scalability, performance, and resiliency of platform data services.
  • Contribute to defining, measuring, and improving service-level indicators, service-level objectives, and error budgets.
  • Develop automation for operational activities such as scaling, failover, backup, recovery, upgrades, and routine maintenance.
  • Build and enhance observability solutions using metrics, logging, tracing, dashboards, and alerts.
  • Troubleshoot issues affecting Cassandra, Aerospike, Kafka/MSK, Redis, and related platform services.
  • Participate in on-call rotations and incident response, contributing to root-cause analysis and the implementation of permanent fixes.
  • Write reliable, maintainable, and well-tested Go code for infrastructure automation, platform services, and operational tooling.
  • Collaborate with engineering, platform, security, and operations teams to integrate and deliver reliable data services.
  • Create and maintain operational documentation, procedures, runbooks, and automation playbooks.
  • Participate in code reviews, technical design discussions, and continuous improvement initiatives.
  • Explore practical applications of AI-assisted automation, anomaly detection, automated remediation, and developer-productivity tooling where appropriate.

Benefits

  • medical
  • dental
  • vision
  • matching 401(k)
  • paid time off
  • wellness program
  • employee discounts for Sony products
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service