The Site Health Platform is a critical part of LinkedIn's Reliability Infrastructure organization, focusing on the end-to-end incident management ecosystem. Our mission is to ensure LinkedIn is always available for members and customers, provide engineers with a more insightful and proactive site-wide reliability ecosystem, and keep business and product owners informed about service disruptions. We manage the full incident lifecycle across thousands of services and multiple regions, including incident response, mitigation, problem management, and post-incident learning. The platforms we develop are fundamental to how LinkedIn detects issues, coordinates incident response, gathers context, and transforms outages and near misses into structured, actionable insights. By converting incidents into data and learnings, we empower teams to systematically enhance reliability over time. Our work influences engineering priorities, infrastructure investments, capacity planning, and executive decision-making, ensuring the network's dependability when it matters most. This role offers exposure to diverse technologies, architectures, and systems hosted in state-of-the-art data centers globally.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior