We’re seeking a future team member for the role of SVP. Site Reliability Engineer to join our Technology team. In this role, you’ll make an impact by designing and implementing end-to-end observability (logs, metrics, traces) across distributed systems, integrating and optimizing tools such as AppDynamics, Dynatrace, Grafana, and Splunk, and developing dashboards, alerts, and telemetry frameworks to provide real-time visibility. You will also identify gaps in monitoring and drive adoption of best practices. Additionally, you will identify repetitive operational work and automate it using code and tooling, build self-healing and auto-remediation solutions, enable scalable, reliable processes through automation and engineering rigor, and improve operational efficiency across production environments. You will troubleshoot and resolve complex production issues across distributed systems, participate in incident management, triage, and root cause analysis, improve monitoring and automation based on recurring incident patterns, and collaborate with support and engineering teams to improve system stability. Furthermore, you will define and measure service health using SLIs/SLOs and key performance metrics, identify system bottlenecks and reliability risks, contribute to performance optimization and capacity planning, and provide input into system architecture to improve resilience and scalability.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Executive
Education Level
No Education Listed