Senior Manager, Data & AI Platforms Reliability

AmeriLifeWashington, NV
$160,000 - $177,000

About The Position

We are seeking a highly technical Senior Manager, Data & AI Platforms Reliability to lead the operational excellence, production reliability, and support of our Enterprise Data & AI Platforms. This leader will own day-to-day production operations across Databricks, enterprise data platforms, Financial Data Platform (FDM), enterprise data products, and associated data services. This role is responsible for ensuring highly available, secure, scalable, and well-governed production platforms while leading production support, incident management, release coordination, platform monitoring, operational readiness, and continuous service improvement. The ideal candidate combines deep technical expertise with strong operational leadership and has experience managing enterprise-scale cloud data platforms.

Requirements

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or related field.
  • 10+ years supporting enterprise data platforms.
  • 5+ years leading technical operations or production support teams.
  • Hands-on Databricks experience, including Delta Lake, Unity Catalog, Spark, Workflows, SQL Warehouses, and Lakehouse architecture.
  • Experience supporting enterprise data warehouses and cloud data platforms.
  • Strong understanding of DevOps, CI/CD, Infrastructure as Code, monitoring, observability, and automation.
  • Experience with Azure cloud services.
  • Strong SQL and performance tuning expertise.
  • Experience leading major incident management and root cause analysis.

Nice To Haves

  • Insurance or Financial Services experience.
  • Experience supporting Financial Data Platforms (FDM).
  • Experience with Data Vault 2.0 and Medallion Architecture.
  • ITIL certification.
  • Experience implementing enterprise monitoring platforms.
  • Experience managing SOX-controlled production environments.

Responsibilities

  • Own production support for enterprise data platforms and data products.
  • Lead incident, problem, change, and release management processes.
  • Ensure platform stability, availability, and operational excellence.
  • Drive root cause analysis and implement permanent corrective actions.
  • Establish operational SLAs, KPIs, and platform health dashboards.
  • Lead production readiness reviews for all new platform releases.
  • Own day-to-day operations of Databricks Lakehouse, FDM, EDR, semantic platforms, and enterprise data services.
  • Partner with Engineering teams to ensure seamless deployments into production.
  • Manage platform capacity planning, performance tuning, resiliency, and operational scalability.
  • Oversee platform upgrades, maintenance windows, and disaster recovery readiness.
  • Establish enterprise monitoring, alerting, logging, and observability standards.
  • Implement automation to reduce operational overhead and improve reliability.
  • Drive continuous improvement initiatives across production operations.
  • Develop operational playbooks, runbooks, and support procedures.
  • Lead Production Support Engineers, Platform Engineers, Database Administrators, and Operations specialists.
  • Build a high-performing operations organization focused on customer experience and reliability.
  • Partner with Data Engineering, Architecture, Governance, Infrastructure, Security, and business stakeholders.

Benefits

  • PTO
  • medical
  • dental
  • vision
  • retirement savings
  • disability insurance
  • life insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service