Lead Site Reliability Engineer

Sherwin-Williams•Cleveland, OH
•Onsite

About The Position

The Lead Site Reliability Engineer role is responsible for optimizing the organization's IT products, services, systems, and digital products. The incumbent works to develop customized solutions to automate the administration, monitoring, and operation of business-critical applications and services, supervises tests for resiliency, redundancy, and failover to ensure uptime, and troubleshoots potential issues to ensure that IT products, services, systems, and digital products are running efficiently and effectively. In addition, the role is responsible for leading the design and implementation of scalable and reliable application and service solutions that can run across multiple environments and technologies. The incumbent fosters a culture of collaboration between cross-functional departments, including development teams, infrastructure teams, and support organizations, to enhance and improve system operability and provide training and knowledge transfer to team members. The role is also responsible for implementing leading practices and emerging technologies that will drive additional efficiencies across the IT organization. The role is also responsible for overseeing project planning, cost analysis, and vendor comparisons when assessing potential solutions and implementing technology and operational improvements.

Requirements

  • Bachelor’s degree in Computer Science or Information Systems, or in lieu of a degree, at least 9 years of experience in the field of site reliability engineering
  • 6+ years of experience in Site Reliability Engineering, Software Engineering, DevOps Engineering, Platform Engineering, Systems Engineering, or a related technical discipline.
  • Experience supporting and operating production applications, services, or enterprise technology platforms.
  • Experience with monitoring, logging, observability, and incident response practices.
  • Experience utilizing automation, scripting, or tooling to improve reliability, operational efficiency, and supportability.
  • Experience collaborating with cross-functional teams to troubleshoot, resolve, and prevent production issues.
  • Strong analytical and problem-solving skills
  • Must be at least (18) eighteen years of age
  • Must be legally authorized to work in the country of employment without now or in the future requiring sponsorship for employment visa status (e.g., OPT, CPT, H1B, EB-1, etc.)

Nice To Haves

  • Experience supporting customer-facing or business-critical digital products and services at a global level.
  • Experience establishing and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • Experience leading major incident response, root cause analysis, and post-incident reviews.
  • Experience with cloud-native architecture (Microsoft Azure, AWS, Kubernetes, etc.)
  • Experience implementing observability, alerting, and operational excellence practices.
  • Experience with CI/CD pipelines and Infrastructure as Code (IaC).
  • Experience driving reliability, resiliency, and performance improvements in distributed systems.
  • Experience with leadership and executive-level communication.
  • Microsoft Certified: Azure Solutions Architect Expert
  • Relevant Site Reliability Engineering, Azure, Cloud, DevOps, or Platform Engineering certifications

Responsibilities

  • Optimize IT products, services, systems, and digital products by proactively analyzing application, service, and operational health metrics to identify potential inefficiencies and thereby ensuring improvements in performance, availability, and reliability.
  • Develop customized solutions to automate the deployment, administration, monitoring, and operation of applications and services and train other team members on the use of automation tools.
  • Develop a formal process for continuously reviewing and monitoring system SLIs, SLOs, SLAs, and OKRs and build optimization plans to address areas of improvement.
  • Foster a culture of collaboration between cross-functional teams to ensure improvement in IT products, services, systems, and digital products.
  • Optimize the use of applications, services, integrations, observability tools, infrastructure components, and load-balancing technologies by tracking performance, identifying potential issues, and ensuring optimal operation.
  • Lead efforts to continuously improve application and service reliability, performance, resiliency, observability, and operational readiness and take steps to mitigate potential issues.
  • Evaluate application and service requirements, lead cross-functional implementation teams, and conduct post-implementation reviews to share lessons learned from the project.
  • Create detailed implementation plans for the integration of new technologies, products, and services into the existing environment that will improve application and service resilience, performance, and reduce costs.
  • Provide leadership and knowledge-sharing to mentor engineers on reliability engineering, observability, automation, incident management, and operational excellence practices and assess their effectiveness.

Benefits

  • Life … with rewards, benefits and the flexibility to enhance your health and well-being
  • Career … with opportunities to learn, develop new skills and grow your contribution
  • Connection … with an inclusive team and commitment to our own and broader communities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service