Senior Data Ops Engineer, Data Activation & Products - Activision

Activision BlizzardSanta Monica, CA
$102,800 - $190,204Hybrid

About The Position

We are looking for a Site Reliability Engineer to help improve the reliability, observability, and operational maturity of our data platforms, Kubernetes-based deployment systems, internal applications, and cloud environments. This role sits at the intersection of SRE, DevOps, and data. The ideal candidate is comfortable operating production systems, troubleshooting across infrastructure and applications, and helping teams deploy and support services more safely. The role does not require someone to be a data engineer, but they should be excited about the systems that support modern data engineering, including Databricks, Spark, Airflow/Astronomer, streaming pipelines, event systems, internal tools, and backend services. We are especially interested in a creative, curious engineer who enjoys learning new technology and using AI-accelerated development practices to solve problems faster and more thoughtfully. You should be excited to experiment with agentic development tools, automation frameworks, and emerging platform capabilities, while applying sound engineering judgment.

Requirements

  • 5+ years of experience in SRE, DevOps, cloud infrastructure, platform engineering, software engineering, data platform operations, or related production-support roles.
  • Hands-on experience supporting Kubernetes-based workloads, deployment systems, cloud infrastructure, or production application environments.
  • Familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git-based workflows, and scripting in Bash, Python, or similar languages.
  • Experience with monitoring, logs, metrics, dashboards, alerting, and incident management practices.
  • Comfort working with event systems such as Kafka, Google Pub/Sub, Kinesis, or similar technologies, including topics, subscriptions, consumers, retries, lag, and dead-letter queues.
  • Strong troubleshooting mindset, clear communication, and comfort operating in a production-support environment.
  • Interest in data platforms, data engineering systems, orchestration, streaming workloads, event-driven architecture, internal developer tools, and production data services.

Nice To Haves

  • Experience with Kubernetes deployment and release tooling such as Helm, ArgoCD, or similar GitOps workflows.
  • Experience with CI/CD automation using GitHub Actions, GitLab CI, Jenkins, or similar pipelines.
  • Familiarity with infrastructure as code using Terraform or similar provisioning tools.
  • Familiarity with Databricks, Spark, Spark Structured Streaming, Airflow, Astronomer, dbt, Kafka, Pub/Sub, object storage, or lakehouse architectures.
  • Experience with observability tools such as Grafana, Prometheus, Cloud Monitoring, Datadog, Splunk, or similar platforms.
  • Experience supporting internal applications, APIs, event-driven services, streaming consumers, or backend workers.
  • Exposure to data reliability concepts such as freshness, latency, completeness, pipeline health, data quality checks, and dependency-aware alerting.
  • Exposure to SLOs, SLAs, error budgets, postmortems, or formal reliability practices.

Responsibilities

  • Monitor service health, respond to alerts, and participate in incident response for cloud, Kubernetes, application, and data platform environments.
  • Investigate reliability issues across Kubernetes, networking, DNS, application runtime behavior, Databricks jobs, Spark workloads, orchestration systems, event systems, and dependent services.
  • Support the reliability of internal applications, APIs, workers, streaming consumers, event-driven services, and deployment workflows used by data engineering and business teams.
  • Build and maintain dashboards, alerting, runbooks, and operational documentation that improve detection and recovery speed.
  • Improve observability for Databricks environments, including job health, Spark streaming workloads, structured streaming metrics, cluster behavior, failures, latency, throughput, and cost signals.
  • Help route Spark streaming metrics, operational logs, event-system signals, and platform health signals into monitoring tools such as Grafana.
  • Contribute to alerting patterns for Databricks workflows, Airflow/Astronomer DAGs, dbt jobs, data freshness, pipeline failures, event lag, dead-letter queues, and production data dependencies.
  • Contribute scripts and automation that reduce repetitive operational work and improve environment hygiene.
  • Support release and deployment reliability by validating changes, improving rollback readiness, and strengthening change safety.
  • Partner with data engineers, analytics engineers, and software engineers to improve reliability across pipelines, services, internal tools, event systems, and data products.
  • Participate in post-incident follow-up and help close corrective actions that prevent recurrence.
  • Support platform modernization and migration efforts, including orchestration platform changes, deployment system improvements, and shared reliability standards.

Benefits

  • Medical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance
  • 401(k) with Company match, tuition reimbursement, charitable donation matching
  • Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leave
  • Mental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, ID protection, rental insurance, and others
  • Relocation assistance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service