Lead Site Reliability Engineer

Wells Fargo BankCharlotte, NC
Onsite

About The Position

Wells Fargo is seeking a Lead Site Reliability Engineer within the Consumer Technology (CT) organization, providing technology solutions to the LOBs and managing the application portfolios through the enablers of skills, stability, security, scalability, speed, and success. Learn more about the career areas and business divisions at wellsfargojobs.com. We are seeking a highly motivated Lead Engineer to drive operational excellence, reliability, and support readiness across a portfolio of data products and platforms. This role combines Site Reliability Engineering (SRE) practices and cross-functional collaboration to ensure stable, resilient, and supportable solutions.

Requirements

  • 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education.
  • 4+ years of strong experience in production support, SRE, or platform engineering in enterprise environments.
  • 3+ years of hands-on experience with observability platforms such as Grafana, Splunk, AppDynamics, or similar monitoring and analytics tools.

Nice To Haves

  • Knowledge of Google Cloud Platform (GCP) services, including Airflow (Cloud Composer), BigQuery, and cloud-native operational practices.
  • Strong understanding of reliability engineering concepts, service-level objectives (SLOs), monitoring strategies, and production support processes.
  • Strong collaboration with cross-functional teams (AppDev, Infra, Business stakeholders).
  • Experience supporting data pipelines and event-driven architectures.
  • Google Cloud Professional Data Engineer Certification.
  • Experience supporting Airflow data pipelines.
  • Familiarity with CI/CD, Infrastructure as Code, and cloud-native operational practices.
  • Proven ability to lead high-severity incidents with effective stakeholder communication and executive reporting.
  • Strong expertise in root cause analysis, problem management, and operational excellence practices.
  • Strong understanding of reliability engineering concepts, service-level objectives (SLOs), monitoring strategies, and production support processes.
  • Strong collaboration with cross-functional teams (AppDev, Infra, Business stakeholders).
  • Demonstrate strong ownership, accountability, and a customer-first mindset.
  • Continuously identify opportunities to improve reliability, supportability, and operational efficiency.
  • Drive outcomes with urgency in a fast-paced, mission-critical environment.
  • Influence cross-functional teams across platform, infrastructure and data engineering.
  • Provide technical leadership, mentorship, and coaching.

Responsibilities

  • Lead production operations across multiple data products.
  • Act as technical escalation lead for critical incidents, driving triage, resolution, and RCA closure.
  • Drive observability maturity across applications (metrics, logging, alerting, dashboards).
  • Ensure runbook quality, support readiness, and operational documentation completeness.
  • Partner with AppDev teams to enforce supportability, reliability, and release readiness standards.
  • Lead problem management including trend analysis and permanent fix tracking.
  • Govern change management ensuring risk-aware deployments and production stability.
  • Mentor engineers and drive cross-training to improve team coverage and resiliency.
  • Participate in on-call rotations and provide leadership during major production incidents.

Benefits

  • Relocation assistance is not available for this position.
  • Visa Sponsorship is not available for this position.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service