We’re seeking a future team member for the role of SVP. Site Reliability Engineer to join our Technology team. This role is located in Lake Mary, FL and Pittsburgh, PA. In this role, you’ll make an impact in the following ways: Design and implement end-to-end observability (logs, metrics, traces) across distributed systems, Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk, Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility, Identify gaps in monitoring and drive adoption of best practices, Identify repetitive operational work and automate it using code and tooling, Build self-healing and auto-remediation solutions, Enable scalable, reliable processes through automation and engineering rigor, Improve operational efficiency across production environments, Troubleshoot and resolve complex production issues across distributed systems, Participate in incident management, triage, and root cause analysis, Improve monitoring and automation based on recurring incident patterns, Collaborate with support and engineering teams to improve system stability, Define and measure service health using SLIs/SLOs and key performance metrics, Identify system bottlenecks and reliability risks, Contribute to performance optimization and capacity planning, Provide input into system architecture to improve resilience and scalability.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed