Senior Systems Operations Engineer

Wells FargoMinneapolis, MN
$42 - $74Hybrid

About The Position

Wells Fargo is seeking a Site Reliability Engineer that combines technical expertise with a passion for automation, resilience, and operational excellence. This role will partner with engineering, and application teams to design, build, and support highly available, scalable, and secure technology platforms while driving reliability improvements through observability, automation, and continuous service optimization. The ideal candidate will possess strong problem-solving skills, a customer-focused mindset, and the ability to lead incident response, root cause analysis, and reliability initiatives that reduce risk and improve the overall stability of enterprise services.

Requirements

  • 4+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 3+ years of demonstrated experience leading Production Application Support at scale, including ITIL-aligned incident, problem, and change management
  • 3+ years of experience implementing application Observability through tools such as AppDynamics, Splunk, Elastic, BigPanda (AIOPS), Grafana, Microsoft Application Insights
  • 3+ years of experience supporting large‑scale Java and .NET applications; strong RDBMS expertise (Oracle, MSSQL); NoSQL experience (MongoDB) a plus

Nice To Haves

  • Experience supporting AI/ML-enabled production systems, including reliability, scalability, and operational risk management for LLM- or agent-based workflows
  • Understanding of AI operational concerns such as model observability, prompt/configuration change management, data drift, failure modes, and integration of AI signals into incident response and SRE practices.
  • Thorough understanding of application environment implementations, including on prem, client server, on prem cloud, hybrid cloud and public cloud
  • Experience using and configuring Continuous Integration Continuous Deploy (CICD) tools including Jenkins, Artifactory, Harness as well as Terraform
  • Proven experience reducing operational toil through automation (Ansible or equivalent), including runbook automation and self‑healing patterns
  • Experience operating resilient systems, including traffic management (F5, AVI), fault tolerance, and data replication

Responsibilities

  • Lead or participate in managing all installed systems and infrastructure within the Systems Operations functional area
  • Contribute in increasing system efficiencies and lowering the human intervention time on related tasks
  • Review and analyze moderately complex operational support systems, application software, and system management tools to ensure the highest levels of systems and infrastructure availability
  • Work with vendors and other technical personnel for problem resolution
  • Lead team to meet technical deliverables while leveraging solid understanding of technical process controls or standards
  • Collaborate with vendors and other technical personnel to resolve technical issues and achieve highest levels of systems and infrastructure availability

Benefits

  • Health benefits
  • 401(k) Plan
  • Paid time off
  • Disability benefits
  • Life insurance, critical illness insurance, and accident insurance
  • Parental leave
  • Critical caregiving leave
  • Discounts and savings
  • Commuter benefits
  • Tuition reimbursement
  • Scholarships for dependent children
  • Adoption reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service