About The Position

This role involves owning and operating a suite of platform services critical for FinOps analytics at ServiceNow. The Senior Software Engineering Manager will lead a team responsible for the health, reliability, and evolution of these services, including distributed query engines, BI platforms, and developer environments. The position requires a blend of technical expertise in open-source data infrastructure, people leadership, and stakeholder management, with a focus on automation, operational excellence, and leveraging AI/ML for platform enhancement. The goal is to ensure these platform services meet stringent SLOs, support migration efforts, and enable users to work efficiently without platform instability.

Requirements

  • Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
  • 12+ years of experience in software or platform engineering, with 5+ years in engineering management leading teams that own production platform services, with a Bachelor’s degree; or 10 years and a Master’s degree; or a PhD with 7 years of experience in Computer Science, Engineering, or a related technical field; or equivalent experience.
  • Proven track record managing teams that operate and scale open-source data infrastructure (query engines, BI platforms, developer environments, or similar) in production.
  • Hands-on experience operating distributed query engines (Trino, Presto, Spark, or similar) including cluster tuning, scaling, and performance optimization.
  • Strong knowledge of Kubernetes and containerized service deployment, enough to architect solutions and debug issues even if a separate team owns the clusters.
  • Demonstrated ability to establish SLOs, observability, and incident-response practices for platform services and to drive operational maturity over time.
  • Experience managing platform upgrades, migrations, and version lifecycle for open-source technologies in production without disrupting users.
  • Proven people leadership. Experience hiring, developing, and retaining strong platform engineers, and building team culture around operational excellence and automation.
  • Strong bias toward automation over manual toil, with experience building or directing the development of internal tooling and self-service workflows.
  • Excellent collaboration skills across engineering, data, DevOps, and business stakeholders.
  • Full professional proficiency in English.
  • Distributed query engines. Trino or Presto operations including deployment, scaling, resource group management, query optimization, connector configuration, and upgrades.
  • Data catalog and lakehouse. Hive Metastore operations and familiarity with modern catalog alternatives (Nessie, AWS Glue, Unity Catalog, Polaris).
  • Apache Iceberg table format concepts.
  • BI and analytics platforms. Operating self-hosted BI tools such as Lightdash, Redash, Metabase, or Superset, including deployment, scaling, SSO integration, and user management.
  • Developer platforms. Coder, JupyterHub, or similar cloud development environment platforms, including workspace provisioning, template management, and resource policies.
  • Observability. Monitoring, alerting, and logging for platform services (Splunk, Prometheus, Grafana, CloudWatch, or similar). SLO design and tracking.
  • Security and access control. SSO/OIDC integration, RBAC, row-level security, secrets management, and audit logging across platform services.
  • Infrastructure familiarity. Kubernetes, Helm, Docker, Infrastructure as Code (Terraform, CDK), and CI/CD pipelines. Enough depth to partner effectively with infra teams and architect platform deployments.
  • Scripting and automation. Python, Bash, or Go for operational tooling, automation, and integration work.
  • Proven ability to balance hands-on technical work with people leadership, knowing when to go deep and when to delegate.
  • Strong technical judgment with the ability to evaluate open-source technologies, make build-vs-buy decisions, and sequence a platform roadmap.
  • Effective stakeholder management across technical and non-technical audiences, translating platform capabilities and constraints into business terms.
  • Strong technical writing and documentation skills for runbooks, architecture decisions, and team processes.
  • Track record of building high-trust, high-ownership engineering teams.

Nice To Haves

  • Direct experience operating Lightdash or dbt-integrated BI platforms.
  • Experience with Project Nessie or other versioned/transactional catalog systems.
  • Experience operating Coder or similar remote development environment platforms at scale.
  • Background in FinOps, cloud cost management, or financial data platforms.
  • Experience with Apache Iceberg table maintenance (compaction, snapshot expiry, partition evolution).
  • Experience in regulated or multi-environment cloud deployments (FedRAMP, GovCloud, or similar).
  • Open-source contributions to data infrastructure tooling.

Responsibilities

  • Own the operational health and reliability of Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore, and Nessie across development and production environments.
  • Establish and maintain SLOs for platform availability, query performance, and workspace provisioning, including building dashboards and alerts.
  • Manage Trino cluster operations end-to-end, including deployment, scaling, upgrades, performance tuning, resource group management, query optimization support, and user access controls.
  • Drive the platform upgrade and patching cadence, balancing stability with security fixes and feature releases.
  • Build runbooks, on-call processes, and incident-response practices for quick resolution and learning from production issues.
  • Ensure platform security across all services, including access controls, authentication (SSO/OIDC integration), secrets management, and audit logging.
  • Lead the migration from Hive Metastore to Nessie as the versioned Iceberg catalog.
  • Drive Lightdash platform improvements including version upgrades, performance optimization, row-level security configuration, and governed self-service analytics.
  • Evolve the Coder platform through workspace template lifecycle management, resource policies, idle-stop tuning, and onboarding new users and use cases.
  • Own the Jupyter and Redash platforms, ensuring availability, scaling, integration with Trino and the lakehouse, and user lifecycle management.
  • Evaluate and adopt new open-source technologies to enhance the platform or reduce operational burden.
  • Manage, mentor, and grow a team of 3 to 5 platform engineers, setting expectations, providing feedback, and creating career development paths.
  • Hire and build the team to match the platform’s growing scope and user base.
  • Foster a culture of operational excellence, automation over toil, and blameless incident retrospectives.
  • Set engineering standards for how the team builds, deploys, monitors, and documents platform services.
  • Partner with the DevOps/infrastructure team on Kubernetes capacity, networking, storage, and CI/CD pipeline needs.
  • Serve as the platform liaison to data engineers, analysts, and FinOps practitioners, understanding their workflows and prioritizing improvements.
  • Collaborate with the Data Platform and Data Governance teams to ensure platform services align with enterprise standards.
  • Support the broader Cloudera-to-lakehouse migration by ensuring Trino, Nessie, and the catalog layer are production-ready.
  • Apply AI/ML tooling where it accelerates platform operations, monitoring, or user support.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service