Senior Production Operations Engineer

PlayStationSan Diego, CA
Hybrid

About The Position

As a member of the Production Operations Engineering team within the platform technology group, you will help keep key user experiences on the platform available, resilient and high performing across time zones and critical business periods, while continually enabling our service teams to deliver new and exciting products and technical features. The team is trusted to respond when production services need support, including through rotational on-call, incident response and urgent operational escalations. You will be empowered to drive and lead technical initiatives, helping identify and proactively drive improvements in process, technology and production operations supporting millions of users. This senior individual contributor role operates at P4 scope: highly independent, hands-on, influential across partner teams, and accountable for meaningful production operations capabilities. The ideal candidate brings deep production experience, strong automation instincts, modern cloud/container operations expertise, and curiosity for using AI-assisted engineering workflows to improve operational quality and delivery speed.

Requirements

  • Senior hands-on engineer equally comfortable with software development, systems engineering and production operations.
  • Proven ability to operate distributed services at scale while automating toil, improving processes and raising production readiness standards.
  • Strong troubleshooting depth across user experience, application runtime, Linux, infrastructure, cloud, Kubernetes/EKS/UKS and network layers.
  • Highly independent technical leader who uses sound judgment, mentors less experienced engineers and drives follow-through during incidents.
  • Distributed service operations, Unix/Linux systems internals, networking, TCP/IP, HTTP/HTTPS, DNS, load balancing and API troubleshooting.
  • AWS service operations, including ALB, Route 53, API Gateway, Lambda, RDS, DynamoDB, ElastiCache and Java/API services.
  • Containers and orchestration, including Docker, Kubernetes, EKS/UKS and Fargate.
  • Automation and software development in Python, Go or Java; source control and configuration management using GitHub/Git, Ansible, Chef or similar.
  • Infrastructure as Code and CI/CD using Terraform, CloudFormation, Jenkins, Spinnaker or comparable tooling.
  • Observability, incident management and collaboration tooling such as Datadog, CloudWatch, Splunk, Grafana, BigPanda, Slack, JIRA and ServiceNow.
  • Strong communication skills with the ability to adjust messaging by audience, translate complex technical issues for cross-functional business partners, and articulate both technical detail and customer/business impact.
  • Data reporting, analytics or operational data platforms such as SQL, MySQL, Oracle, Snowflake or big data systems.
  • AI-assisted development or agentic engineering workflow experience using tools such as Codex, Claude, Cursor or similar.
  • BS degree or equivalent in Computer Science, Software Engineering or related technical area, 5+ years operating and supporting services in production environments at scale.
  • 3+ years AWS Cloud experience deploying, tuning and operating Java/API services.
  • Hands-on incident management experience, including on-call, incident coordination, stakeholder communications, RCA and corrective-action tracking.

Nice To Haves

  • Payment processing, commerce, entitlement, wallet, fraud or financial systems experience is a plus.

Responsibilities

  • Own application operations and production support for internal and public-facing services in AWS and container environments, with focus on availability, resiliency, scalability, performance and security.
  • Drive production readiness for new services and features, including provisioning, automation, monitoring, alerting, dashboards, runbooks, rollback planning and operational acceptance criteria.
  • Build scripts, tools and repeatable workflows that reduce toil, improve incident response, and standardize operational practices across the environment.
  • Improve observability and event correlation across the platform using actionable service health signals, dashboards, alerts and incident workflows that reduce MTTD and MTTR.
  • Support and improve incident management practices using tools such as Slack, JIRA, ServiceNow, BigPanda and related escalation/notification systems.
  • Partner with SRE, data services, CI/CD, service engineering, platform hosting and product teams to improve end-to-end reliability and operational readiness.
  • Drive performance, capacity and cost optimization for services using AWS, Kubernetes, EKS/UKS, autoscaling and related cloud patterns.
  • Use AI-assisted and agentic engineering tools, such as Codex, Claude, Cursor or similar, to accelerate development, operational automation and production-signal analysis where appropriate.
  • Provide rotational on-call support, lead incident triage/mitigation, and document post-incident reviews with corrective actions and prevention opportunities.

Benefits

  • medical
  • dental
  • vision
  • matching 401(k)
  • paid time off
  • wellness program
  • employee discounts for Sony products
  • bonus package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service