Staff Software Engineer, Reliability

LinkedInMountain View, CA
$156,000 - $255,000Hybrid

About The Position

The Site Health Platform is a critical part of LinkedIn's Reliability Infrastructure organization, focusing on the end-to-end incident management ecosystem. Our mission is to ensure LinkedIn is always available for members and customers, provide engineers with a more insightful and proactive site-wide reliability ecosystem, and keep business and product owners informed about service disruptions. We manage the full incident lifecycle across thousands of services and multiple regions, including incident response, mitigation, problem management, and post-incident learning. The platforms we develop are fundamental to how LinkedIn detects issues, coordinates incident response, gathers context, and transforms outages and near misses into structured, actionable insights. By converting incidents into data and learnings, we empower teams to systematically enhance reliability over time. Our work influences engineering priorities, infrastructure investments, capacity planning, and executive decision-making, ensuring the network's dependability when it matters most. This role offers exposure to diverse technologies, architectures, and systems hosted in state-of-the-art data centers globally.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or related technical field or equivalent practical experience.
  • 6+ years of professional experience in software development, distributed systems, or reliability engineering.
  • Experience leading technical projects/providing architectural leadership
  • Experience building products and operating large-scale distributed systems.
  • Experience with two or more backend languages such as Go, Python or Java with a track record of owning complex production systems.
  • Full-stack engineering experience, including building user-facing web applications and operational dashboards using modern frontend frameworks such as React.js, along with backend APIs and data pipelines.
  • Understanding of web development fundamentals including API design, performance, accessibility and building intuitive interfaces for engineers and operational users.
  • Understanding of reliability engineering principles, incident management, observability and operating systems under failure conditions.
  • Demonstrated ability to lead technical design across teams, influence architecture beyond direct ownership and drive adoption through well-designed platforms.
  • Debugging and root cause analysis skills, with the ability to communicate complex technical findings clearly to engineers, partners and leadership.

Nice To Haves

  • Experience applying AI or LLM-based techniques to operational or incident data, including automated summarization, classification, root cause hypothesis generation or reliability recommendations.
  • Familiarity with vector databases and retrieval-based systems used to power context-aware analytics, search or agentic workflows.
  • Frontend engineering experience beyond basic UI, including building data-dense, high-signal interfaces for engineers using React.js, modern state management and visualization libraries.
  • Experience designing end-to-end full-stack systems where frontend, backend, data and reliability concerns are considered holistically.
  • Background in building internal developer platforms, observability tools, or incident response systems used at scale.
  • A demonstrated ability to simplify complex workflows, reduce operational toil and replace manual processes with well-designed automation.

Responsibilities

  • Designing and evolving the core incident management platforms that power LinkedIn’s full incident lifecycle, from detection and response to problem management and prevention, across thousands of services and teams.
  • Serving in a critical on-call rotation, providing expert incident triage and coordination during high-severity outages. Partnering closely with service owners and product teams to diagnose issues quickly, mitigate member impact, and drive timely resolution under pressure.
  • Transforming raw, unstructured incident data into clear, actionable intelligence using AI and LLM-based systems, including automated summarization, classification, root cause signals, and mitigation recommendations.
  • Building analytics and insights that surface systemic reliability risks, recurring failure patterns, and cross-service dependencies, enabling org-level prioritization rather than isolated, service-by-service fixes.
  • Building platforms and tools that enable realistic, fleet-wide stress testing of data center and regional capacity, validating incident readiness across dependencies, traffic patterns, and growth scenarios before they impact a significant production outage.
  • Driving consistency, clarity, and quality in how incidents are declared, managed, reviewed, and learned from, raising the reliability bar across a large, fast-moving engineering organization.
  • Influencing service architecture, SLOs, and reliability standards through platforms, data, and technical leadership, ensuring improvements are durable, measurable, and adopted at scale.

Benefits

  • annual performance bonus
  • stock
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service