Reliability Engineer 3 (Observability Specialist)

U.S. Bank National AssociationBrookfield, WI
$98,175 - $115,500Hybrid

About The Position

At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions, enabling the communities we support to grow and succeed in the right ways, all more confidently and more often—that’s what we call the courage to thrive. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive. Try new things, learn new skills and discover what you excel at—all from Day One. As a wholly owned subsidiary of U.S. Bank, Elavon is committed to building the platforms and ecosystems that help over 1.5 million customers around the world to achieve their financial goals—no matter what they need. From transaction processing to customer service, to driving innovation and launching new products, we’re building a range of tailored payment solutions powered by the latest technology. As part of our team, you can explore what motivates and energizes your career goals: partnering with our customers, our communities, and each other.

Requirements

  • Bachelor's degree, or equivalent work experience
  • Five to seven years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development

Nice To Haves

  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering.
  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring.
  • Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products.
  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks.
  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry.
  • Experience building, standardizing, and tuning operational dashboards and actionable alerts that communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage.
  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes.
  • Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements.
  • Excellent stakeholder management, communication, and technical leadership skills.

Responsibilities

  • Own and drive the enterprise Observability Strategy, aligning critical customer journeys, reliability objectives, operational excellence goals, and business outcomes across multiple platforms and technology domains.
  • Define and govern enterprise-wide SLIs, SLOs, Error Budgets, KPIs, and Reliability Standards, establishing measurable service health objectives and accountability frameworks.
  • Architect scalable observability solutions leveraging telemetry, distributed tracing, logging, metrics, synthetic monitoring, and APM/RUM capabilities to proactively improve service reliability and customer experience.
  • Establish and oversee Observability Governance frameworks, including instrumentation standards, telemetry policies, dashboard lifecycle management, alert governance, and monitoring best practices.
  • Serve as a trusted advisor to Product, Engineering, SRE, Infrastructure, and Operations leaders, influencing technology strategy, production readiness, resilience planning, and operational risk management.
  • Lead the design and evolution of executive and operational service health reporting, delivering actionable insights into availability, latency, customer impact, dependency performance, reliability trends, and SLO compliance.
  • Drive continuous improvement through advanced analysis of telemetry, incidents, problem management data, and alert effectiveness, reducing operational noise, accelerating detection, and improving service resiliency.
  • Provide senior technical leadership, mentorship, and standards governance for observability engineering practices, including distributed tracing, logging, metrics, synthetic testing, monitoring architecture, and incident detection strategies across the enterprise.

Benefits

  • Healthcare (medical, dental, vision)
  • Basic term and optional term life insurance
  • Short-term and long-term disability
  • Pregnancy disability and parental leave
  • 401(k) and employer-funded retirement plan
  • Paid vacation (from two to five weeks depending on salary grade and tenure)
  • Up to 11 paid holiday opportunities
  • Adoption assistance
  • Sick and Safe Leave accruals of one hour for every 30 worked, up to 80 hours per calendar year unless otherwise provided by law
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service