About The Position

Become a member of a team where you can contribute significantly to shaping the future of a world-renowned and influential company. Among top performers, you can make a direct and meaningful impact. As a Senior Lead Infrastructure Engineer at JPMorganChase within the Corporate Sector – Infrastructure Platforms, you exhibit both depth and breadth of knowledge regarding software, applications, and technical processes across multiple technical disciplines. You also have a specialization in a specific domain within infrastructure engineering to drive programs or initiatives consisting of multiple technologies and applications.

Requirements

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Own reliability and operational excellence for IP-based storage services, including availability, performance, capacity, and operational risk reduction.
  • Act as the SME for enterprise storage platforms (e.g., Dell EMC PowerFlex, NetApp SolidFire, Pure Storage, PMAX) including design input, troubleshooting, upgrades, and lifecycle management.
  • Provide 24x7 support coverage, incident response, and escalations; lead restoration activities during high-severity events.
  • Drive SRE best practices; Define, measure, and improve SLIs/SLOs, error budgets, and operational KPIs for storage services.
  • Build and maintain observability: metrics, logs, traces (where applicable), dashboards, alerts, and runbooks.
  • Reduce operational load by identifying and eliminating toil through automation; Automate repetitive tasks (provisioning, health checks, failover validation, reporting, hygiene tasks).
  • Implement safe automation with guardrails, change controls, and rollback strategies.
  • Expert knowledge of AAAS automation.
  • Perform root cause analysis (RCA) and problem management; implement corrective and preventive actions to prevent recurrence.
  • Maintain and continuously improve runbooks, standard operating procedures, on-call playbooks, and knowledge articles.
  • Collaborate with engineering, network, compute, and platform teams on architecture reviews, change planning, and reliability improvements.
  • Support capacity management and performance engineering: forecasting, trending, saturation analysis, and proactive remediation.
  • Ensure compliance with operational standards (change management, risk controls, documentation, audit readiness).
  • Role participates in a 24x7 on-call rotation and is expected to respond to production incidents within defined SLAs.
  • May require off-hours work for planned maintenance, upgrades, and risk-reduction activities.

Nice To Haves

  • Experience with storage networking and adjacent technologies (VLANs, MTU/jumbo frames, latency analysis, QoS).
  • Understanding of resilience patterns: redundancy, failure domains, replication, backup/restore, DR testing.
  • Experience driving operational maturity initiatives (SLO rollout, runbook standardization, automation roadmaps).
  • Clear SLOs/SLIs and dashboards adopted by stakeholders; fewer false positives and more actionable alerts.
  • Improved platform stability and predictable performance/capacity posture.

Responsibilities

  • Applies deep technical expertise and problem-solving methodologies focused on analyzing complex data and systems, anticipating issues, and finding ways to mitigate risk
  • Works with other platforms to architect and implement changes required to resolve issues and modernize the organization and its technology processes
  • Be responsible for infrastructure engineering in accordance with business requirements
  • Executes work according to compliance standards, risk and security, and business objectives
  • Own end-to-end problem detection, resolution, and prevention; lead efforts to improve MTTD, MTTM, and MTTR through better observability, triage, and remediation practices.
  • Drive a workstream or project spanning one or more infrastructure engineering technologies, including technology lifecycle management and dependency management across network, compute, database, and application platforms.
  • Partner with adjacent platform teams to architect and implement changes that resolve systemic issues, reduce operational risk, and modernize technology and operational processes.
  • Design and deliver creative, scalable solutions for high-complexity engineering challenges, including development of automation, tooling, and repeatable operational patterns.
  • Evaluate upstream/downstream impacts across systems, data flows, and integrations; proactively identify risks and provide clear mitigation and contingency recommendations.
  • Expertise to continuously improve telemetry and observability of storage products with tools like Grafana, Dynatrace, Prometheus, Splunk, Netcool etc.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service