Senior Vice President, Site Reliability Engineer

BNY MellonLake Mary, FL
$120,000 - $190,000Onsite

About The Position

We’re seeking a future team member for the role of SVP. Site Reliability Engineer to join our Technology team. In this role, you’ll make an impact by designing and implementing end-to-end observability (logs, metrics, traces) across distributed systems, integrating and optimizing tools such as AppDynamics, Dynatrace, Grafana, and Splunk, and developing dashboards, alerts, and telemetry frameworks to provide real-time visibility. You will also identify gaps in monitoring and drive adoption of best practices. Additionally, you will identify repetitive operational work and automate it using code and tooling, build self-healing and auto-remediation solutions, enable scalable, reliable processes through automation and engineering rigor, and improve operational efficiency across production environments. You will troubleshoot and resolve complex production issues across distributed systems, participate in incident management, triage, and root cause analysis, improve monitoring and automation based on recurring incident patterns, and collaborate with support and engineering teams to improve system stability. Furthermore, you will define and measure service health using SLIs/SLOs and key performance metrics, identify system bottlenecks and reliability risks, contribute to performance optimization and capacity planning, and provide input into system architecture to improve resilience and scalability.

Requirements

  • 9+ years of experience in Site Reliability Engineering, Software Engineering
  • Strong programming background in Java (preferred) or another modern language
  • Experience with at least one observability platform: AppDynamics, Dynatrace, Grafana, or Splunk
  • Hands-on experience supporting and troubleshooting production systems
  • Strong analytical and problem-solving skills
  • Ability to identify inefficiencies and drive automation

Nice To Haves

  • Experience with distributed systems or microservices architectures
  • Familiarity with CI/CD pipelines and DevOps practices
  • Exposure to cloud platforms and/or Kubernetes
  • Experience scripting (Python, Bash, etc.) for automation
  • Knowledge of SRE concepts like observability, incident management, and reliability engineering

Responsibilities

  • Design and implement end-to-end observability (logs, metrics, traces) across distributed systems
  • Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
  • Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
  • Identify gaps in monitoring and drive adoption of best practices
  • Identify repetitive operational work and automate it using code and tooling
  • Build self-healing and auto-remediation solutions
  • Enable scalable, reliable processes through automation and engineering rigor
  • Improve operational efficiency across production environments
  • Troubleshoot and resolve complex production issues across distributed systems
  • Participate in incident management, triage, and root cause analysis
  • Improve monitoring and automation based on recurring incident patterns
  • Collaborate with support and engineering teams to improve system stability
  • Define and measure service health using SLIs/SLOs and key performance metrics
  • Identify system bottlenecks and reliability risks
  • Contribute to performance optimization and capacity planning
  • Provide input into system architecture to improve resilience and scalability

Benefits

  • Highly competitive compensation
  • Benefits and wellbeing programs
  • Access to flexible global resources and tools
  • Generous paid leaves
  • Paid volunteer time
  • 401(k) plan
  • Company-sponsored medical, dental, vision, and basic life insurance plans
  • Various paid time off benefits, such as vacation and sick time
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service