Site Reliability Engineer

OneAZ Credit UnionPhoenix, AZ
$96,664 - $120,830Onsite

About The Position

The Site Reliability Engineer is responsible for ensuring the reliability, availability, performance, and recoverability of OneAZ's technology platforms and infrastructure. This position combines infrastructure engineering, automation, monitoring, and resiliency practices to maintain highly available systems supporting associates and members. The engineer designs and implements automation solutions using PowerShell and other scripting technologies, administers enterprise monitoring platforms, leads disaster recovery testing activities, and partners with technology teams to improve operational resilience, service reliability, and recovery readiness across on-premises and cloud environments. This role serves as a key contributor to incident response, infrastructure modernization, and continuous improvement initiatives focused on reducing operational risk and improving system uptime. The Site Reliability Engineer works closely with Infrastructure, Information Security, Application Support, Enterprise Architecture, and business teams to identify operational risks, strengthen recovery capabilities, and improve the overall resilience of technology services that support associates and members.

Requirements

  • High School Diploma Required
  • Bachelor's Degree in Information Technology, Computer Science, Information Systems, Engineering, or a related technical field; or equivalent combination of education and experience. Required
  • 5-8 years similar or related experience of experience supporting enterprise infrastructure environments and administering infrastructure monitoring platforms. Required
  • Deep knowledge of networking and cloud technologies, Windows Server, virtualization, and storage with hands on experience supporting and optimizing enterprise infrastructure environments.
  • Experience with disaster recovery testing, failover planning, or business continuity initiatives.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Ability to manage multiple priorities and projects in a fast-paced environment.
  • Strong knowledge of infrastructure monitoring and alerting best practices.
  • Understanding of disaster recovery, business continuity, and operational resiliency concepts.
  • Ability to analyze system performance trends and identify potential issues before business impact occurs.
  • Strong technical documentation and organizational skills.
  • Ability to coordinate activities across multiple technology teams.
  • Strong verbal and written communication skills.
  • Ability to remain organized and effective during critical incidents and recovery activities.
  • Commitment to continuous improvement and operational excellence.

Nice To Haves

  • Experience administering SolarWinds in a large enterprise environment. Preferred
  • Experience within a financial institution, credit union, or regulated industry. Preferred
  • Experience supporting hybrid cloud environments including Microsoft Azure. Preferred
  • Experience coordinating disaster recovery exercises and recovery readiness assessments. Preferred
  • Familiarity with ITIL service management principles. Preferred
  • Industry certifications: SolarWinds Certified Professional (SCP), Microsoft Azure Administrator Associate, Microsoft Azure Solutions Architect, Google Professional Site Reliability Engineer, ITIL Foundation, CBCP, or related certifications.

Responsibilities

  • Administer, maintain, and optimize SolarWinds and other enterprise monitoring platforms.
  • Develop and maintain monitoring dashboards, alerts, reports, and performance metrics.
  • Proactively monitor and improve infrastructure health, availability, capacity, and performance, identifying and mitigating potential issues before they impact service reliability or user experience.
  • Proactively identify infrastructure risks, reliability concerns, and opportunities for improvement.
  • Lead disaster recovery planning, testing, failover exercises, and recovery validation activities.
  • Develop and maintain disaster recovery documentation, runbooks, recovery procedures, and test plans.
  • Coordinate and execute system failover and failback activities for critical applications and infrastructure.
  • Partner with Infrastructure, Security, and Application teams to improve system resiliency and operational readiness.
  • Support business continuity planning initiatives and technology recovery efforts.
  • Lead the investigation and resolution of complex infrastructure and service reliability incidents, conduct root cause analyses (RCAs), identify underlying systemic issues, and drive corrective and preventive actions to improve availability, performance, resiliency and operational excellence.
  • Track, document, and report on disaster recovery testing results, remediation activities, and recovery readiness metrics.
  • Maintain infrastructure diagrams, recovery documentation, and operational procedures.
  • Support audits, examinations, and compliance activities related to disaster recovery, resiliency, and infrastructure operations.
  • Research and recommend technologies and best practices that improve reliability, monitoring, and recoverability.
  • Participate in infrastructure maintenance activities, upgrades, and projects.
  • Develop and maintain automation scripts using PowerShell, Python, Bash, or similar tools to streamline infrastructure operations, automate remediation of common issues, and enhance service availability and performance.

Benefits

  • Generous paid time off: paid holidays, floating holidays, personal days, vacation days, plus sick time
  • Low-cost Medical, Dental & Vision plans
  • Paid childcare assistance
  • Award-winning 401K
  • Gym fee reimbursement
  • Tuition Reimbursement
  • Student loan repayment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service