About The Position

In this role, you’ll make an impact in the following ways: Lead day-to-day production services support for applications within the Fund and Investor Solutions platform. Ensure high availability, performance, and operational stability of production systems. Manage and coordinate major incident response, service restoration, escalation handling, and stakeholder communications. Drive root cause analysis and permanent remediation for recurring incidents and platform issues. Partner with engineering and development teams to improve application resiliency, supportability, and release quality. Implement and enhance monitoring, alerting, and observability capabilities across applications and infrastructure. Oversee production readiness, change implementation, release support, and operational governance. Identify automation opportunities to reduce manual effort and improve service efficiency. Track service metrics, incident trends, platform health indicators, and operational risk items. Support capacity planning, environment stability, continuity, recovery readiness, and controls compliance. Mentor production support teams and foster a culture of accountability, collaboration, and operational excellence.

Requirements

  • Strong experience supporting enterprise production applications in a high-availability environment.
  • Deep understanding of incident management, problem management, change management, and production governance.
  • Hands-on knowledge of application monitoring and observability tools such as Splunk, AppDynamics, Dynatrace, Grafana, Prometheus, or similar platforms.
  • Experience with Unix/Linux environments and strong scripting skills in Shell, Python, or Perl for automation and diagnostics.
  • Strong knowledge of SQL and experience troubleshooting across relational databases such as Oracle, SQL Server, PostgreSQL, or similar.
  • Familiarity with middleware and integration technologies, including APIs, messaging systems, and batch processing frameworks.
  • Experience supporting distributed systems in cloud and hybrid environments, including infrastructure, networking, and platform dependencies.
  • Understanding of release management, DevOps practices, CI/CD pipelines, and environment controls.
  • Working knowledge of ITIL-based service management processes.
  • Strong troubleshooting skills across application, infrastructure, database, and network layers.
  • Experience with scheduling tools, job monitoring, and operational workflows in complex enterprise ecosystems.
  • Knowledge of resilience practices including high availability, disaster recovery, capacity planning, and performance tuning.

Responsibilities

  • Lead day-to-day production services support for applications within the Fund and Investor Solutions platform.
  • Ensure high availability, performance, and operational stability of production systems.
  • Manage and coordinate major incident response, service restoration, escalation handling, and stakeholder communications.
  • Drive root cause analysis and permanent remediation for recurring incidents and platform issues.
  • Partner with engineering and development teams to improve application resiliency, supportability, and release quality.
  • Implement and enhance monitoring, alerting, and observability capabilities across applications and infrastructure.
  • Oversee production readiness, change implementation, release support, and operational governance.
  • Identify automation opportunities to reduce manual effort and improve service efficiency.
  • Track service metrics, incident trends, platform health indicators, and operational risk items.
  • Support capacity planning, environment stability, continuity, recovery readiness, and controls compliance.
  • Mentor production support teams and foster a culture of accountability, collaboration, and operational excellence.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service