Senior Software Engineering Manager - FinOps Platform Services

ServiceNowPleasanton, CA
$190,900 - $334,100Hybrid

About The Position

This role involves owning and evolving platform services critical for FinOps analytics at ServiceNow. The successful candidate will lead a team of platform engineers, focusing on operational health, reliability, and roadmap development for open-source technologies like Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore, and Nessie. The position emphasizes a culture of operational excellence, automation, and collaboration with various engineering and business teams. The goal is to ensure these platform services meet high standards of performance, security, and availability, supporting ServiceNow's global cloud spend analytics.

Requirements

  • Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
  • 12+ years of experience in software or platform engineering, with 5+ years in engineering management leading teams that own production platform services, with a Bachelor’s degree; or 10 years and a Master’s degree; or a PhD with 7 years of experience in Computer Science, Engineering, or a related technical field; or equivalent experience.
  • Proven track record managing teams that operate and scale open-source data infrastructure (query engines, BI platforms, developer environments, or similar) in production.
  • Hands-on experience operating distributed query engines (Trino, Presto, Spark, or similar) including cluster tuning, scaling, and performance optimization.
  • Strong knowledge of Kubernetes and containerized service deployment, enough to architect solutions and debug issues even if a separate team owns the clusters.
  • Demonstrated ability to establish SLOs, observability, and incident-response practices for platform services and to drive operational maturity over time.
  • Experience managing platform upgrades, migrations, and version lifecycle for open-source technologies in production without disrupting users.
  • Proven people leadership. Experience hiring, developing, and retaining strong platform engineers, and building team culture around operational excellence and automation.
  • Strong bias toward automation over manual toil, with experience building or directing the development of internal tooling and self-service workflows.
  • Excellent collaboration skills across engineering, data, DevOps, and business stakeholders.
  • Full professional proficiency in English.
  • Distributed query engines. Trino or Presto operations including deployment, scaling, resource group management, query optimization, connector configuration, and upgrades.
  • Data catalog and lakehouse. Hive Metastore operations and familiarity with modern catalog alternatives (Nessie, AWS Glue, Unity Catalog, Polaris). Apache Iceberg table format concepts.
  • BI and analytics platforms. Operating self-hosted BI tools such as Lightdash, Redash, Metabase, or Superset, including deployment, scaling, SSO integration, and user management.
  • Developer platforms. Coder, JupyterHub, or similar cloud development environment platforms, including workspace provisioning, template management, and resource policies.
  • Observability. Monitoring, alerting, and logging for platform services (Splunk, Prometheus, Grafana, CloudWatch, or similar). SLO design and tracking.
  • Security and access control. SSO/OIDC integration, RBAC, row-level security, secrets management, and audit logging across platform services.
  • Infrastructure familiarity. Kubernetes, Helm, Docker, Infrastructure as Code (Terraform, CDK), and CI/CD pipelines. Enough depth to partner effectively with infra teams and architect platform deployments.
  • Scripting and automation. Python, Bash, or Go for operational tooling, automation, and integration work.
  • Proven ability to balance hands-on technical work with people leadership, knowing when to go deep and when to delegate.
  • Strong technical judgment with the ability to evaluate open-source technologies, make build-vs-buy decisions, and sequence a platform roadmap.
  • Effective stakeholder management across technical and non-technical audiences, translating platform capabilities and constraints into business terms.
  • Strong technical writing and documentation skills for runbooks, architecture decisions, and team processes.
  • Track record of building high-trust, high-ownership engineering teams.

Nice To Haves

  • Direct experience operating Lightdash or dbt-integrated BI platforms.
  • Experience with Project Nessie or other versioned/transactional catalog systems.
  • Experience operating Coder or similar remote development environment platforms at scale.
  • Background in FinOps, cloud cost management, or financial data platforms.
  • Experience with Apache Iceberg table maintenance (compaction, snapshot expiry, partition evolution).
  • Experience in regulated or multi-environment cloud deployments (FedRAMP, GovCloud, or similar).
  • Open-source contributions to data infrastructure tooling.

Responsibilities

  • Own the operational health and reliability of Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore, and Nessie across development and production environments.
  • Establish and maintain SLOs for platform availability, query performance, and workspace provisioning, including building dashboards and alerting.
  • Own Trino cluster operations end-to-end, including deployment, scaling, upgrades, performance tuning, resource group management, query optimization support, and user access controls.
  • Drive the platform upgrade and patching cadence, balancing stability with security fixes and feature releases.
  • Build runbooks, on-call processes, and incident-response practices for quick resolution of production issues.
  • Ensure platform security across all services, including access controls, authentication (SSO/OIDC integration), secrets management, and audit logging.
  • Lead the migration from Hive Metastore to Nessie as the versioned Iceberg catalog.
  • Drive Lightdash platform improvements including version upgrades, performance optimization, row-level security configuration, and governed self-service analytics.
  • Evolve the Coder platform through workspace template lifecycle management, resource policies, idle-stop tuning, and onboarding new users and use cases.
  • Own the Jupyter and Redash platforms, ensuring availability, scaling, integration with Trino and the lakehouse, and user lifecycle management.
  • Evaluate and adopt new open-source technologies to enhance platform capabilities or reduce operational burden.
  • Manage, mentor, and grow a team of 3 to 5 platform engineers, setting expectations, providing feedback, and creating career development paths.
  • Hire and build the team to match the platform’s growing scope and user base.
  • Foster a culture of operational excellence, automation over toil, and blameless incident retrospectives.
  • Set engineering standards for how the team builds, deploys, monitors, and documents platform services.
  • Partner with the DevOps/infrastructure team on Kubernetes capacity, networking, storage, and CI/CD pipeline needs.
  • Serve as the platform liaison to data engineers, analysts, and FinOps practitioners, understanding their workflows and prioritizing improvements.
  • Collaborate with Data Platform and Data Governance teams to ensure platform services align with enterprise standards.
  • Support the broader Cloudera-to-lakehouse migration.
  • Apply AI/ML tooling where it accelerates platform operations, monitoring, or user support.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service