Lead Site Reliability Engineer

JPMorganChase•Jersey City, NJ

About The Position

As a Lead Site Reliability Engineer at JPMorgan Chase within the Asset and Wealth Management, Tech Production and Infrastructure Delivery team, you will be responsible for improving reliability, resilience, and operational performance across a hybrid technology environment spanning modern distributed platforms and mainframe systems.

Requirements

  • Relevant experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or a similar reliability-focused role, including leadership of technical initiatives.
  • Strong knowledge of operating and supporting distributed systems in production (e.g., Linux, networking, middleware, containers and/or cloud platforms).
  • Hands-on experience with monitoring/observability platforms and practices (metrics, logs, traces), including dashboarding and alert engineering.
  • Demonstrated ability to automate operational workflows using one or more scripting/programming languages (e.g., Python, Go, Shell) and standard automation approaches (CI/CD, infrastructure-as-code).
  • Experience supporting or integrating mainframe systems into enterprise operations (monitoring, incident response, operational processes).
  • Strong communication and stakeholder management skills, with ability to lead cross-team reliability improvements.

Nice To Haves

  • Experience establishing SLIs/SLOs and reliability reporting at service or platform level.
  • Familiarity with ITSM/incident tooling, on-call operations, and operational maturity improvements.
  • Experience with resilience patterns (graceful degradation, failover, rate limiting) and reliability testing (chaos testing, load/performance testing).
  • Exposure to regulated or high-control environments and operational risk management practices.

Responsibilities

  • Lead adoption and operationalization of SRE practices, including SLIs/SLOs, error budgets, reliability reviews, and blameless post-incident processes.
  • Design, implement, and continuously improve monitoring and observability capabilities across metrics, logs, traces, and event telemetry to support faster detection and diagnosis.
  • Establish actionable alerting standards, dashboards, and runbooks to improve operational readiness and reduce noise.
  • Drive automation initiatives (self-service, self-healing, automated remediation, CI/CD operational controls, and standardized tooling) to reduce manual effort and improve consistency.
  • Identify, measure, and reduce operational toil through process optimization, tooling enhancements, and platform improvements.
  • Improve incident management practices, including incident response coordination, escalation paths, and continuous improvement based on root cause analysis.
  • Support capacity planning, performance engineering, and resilience testing to strengthen availability and service stability.
  • Partner with application, infrastructure, and operations teams across distributed and mainframe domains to standardize reliability patterns and operational controls.
  • Contribute to governance and operational excellence, including documentation, control evidence where applicable, and operational health reporting.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service