Principal Staff Software Engineer, Systems Infrastructure

LinkedInMountain View, CA
Remote

About The Position

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. This role may be remote or hybrid. At LinkedIn, hybrid roles are performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Remote roles are performed from the designated home work location upon time of hire, and any changes to this home work location requires a review of remote status and approval. LinkedIn’s Reliability Infrastructure team is responsible for defining and driving the reliability strategy, standards, and practices that keep LinkedIn’s most critical systems stable, resilient, and available at massive scale. As a Principal Staff Software Engineer, Reliability Infrastructure, you will serve as a senior technical authority for reliability across LinkedIn Engineering. You will help define how critical services are designed, built, operated, and measured, partnering broadly across infrastructure and product engineering teams to improve resiliency, reduce incidents, and raise the reliability bar across the company. A key focus of this role is driving the adoption and evolution of LinkedIn’s service criticality framework, including reliability expectations for the most business-critical systems. You will help classify services based on criticality and blast radius, define appropriate reliability standards, and influence system architecture to ensure the right levels of availability, redundancy, observability, and failure handling are in place. As AI-assisted software development, agent-based automation, and autonomous operational systems become more prevalent, this role will also help define how LinkedIn safely builds and operates reliable AI-enabled systems. You will shape standards for evaluating, deploying, monitoring, and governing AI-generated code and agentic workflows, ensuring that automation introduced into critical environments is observable, explainable, auditable, and designed with appropriate safeguards, rollback mechanisms, and human oversight. This is not a traditional SRE role focused on operating a single service or team. It is a company-wide technical leadership role for someone with deep distributed systems expertise, strong reliability judgment, and the ability to influence architecture and engineering practices across large organizations.

Requirements

  • BA/BS degree in Computer Science or related technical field, or equivalent practical experience
  • 10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or reliability engineering
  • 5+ years of experience in a technical leadership, architect, or principal-level engineering role
  • Experience designing, building, or operating large-scale distributed systems
  • Experience defining or driving reliability standards such as SLOs, SLIs, uptime targets, incident reduction, or operational readiness frameworks
  • Understanding of high availability, redundancy, fault tolerance, failure modes, and resiliency patterns
  • Experience influencing architecture and engineering practices across multiple teams or organizations
  • Software engineering experience in one or more languages such as Java, Go, C++, Python, or similar

Nice To Haves

  • MS or PhD in Computer Science or related technical field
  • Experience operating at company-wide or large org-wide scope as a reliability, infrastructure, SRE, or production engineering technical leader
  • Experience with tiered service criticality models, priority-based reliability frameworks, or large-scale reliability governance
  • Deep expertise in distributed systems reliability, service resilience, and failure isolation at scale
  • Experience with incident management, postmortems, operational reviews, and driving long-term corrective actions across organizations
  • Experience with observability, monitoring, alerting, capacity planning, disaster recovery, and multi-region failover strategies
  • Background in mature SRE, production engineering, platform reliability, or infrastructure resilience environments
  • Experience driving reliability transformations across large engineering organizations
  • Executive-level communication skills with the ability to align technical decisions to business impact
  • Demonstrated ability to influence technical direction without direct authority and drive adoption of standards across teams
  • Prior work on self-healing or auto-remediation systems at companies with large-scale infrastructure (hyperscalers, large internet companies).
  • Experience defining reliability, safety, or governance standards for AI-enabled systems, agentic workflows, or AI-assisted software development
  • Familiarity with LLM and agent evaluation, production monitoring, guardrails, human oversight, and rollback strategies for autonomous systems
  • Familiarity with emerging standards and frameworks for agentic AI safety and evaluation (evals pipelines, red-teaming autonomous systems, policy guardrails for production agents).

Responsibilities

  • Define and drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering
  • Lead adoption and evolution of service criticality models that set reliability expectations based on business impact and blast radius
  • Serve as a technical authority for architecture decisions related to reliability, resiliency, availability, and failure handling
  • Partner with infrastructure and product engineering teams to improve system design, reduce incident risk, and strengthen operational readiness
  • Identify high-risk systems and drive cross-organizational initiatives to improve reliability of critical services
  • Establish and evolve reliability standards including SLOs, SLIs, uptime expectations, redundancy, monitoring, alerting, and failover patterns
  • Influence engineering culture by promoting reliability-focused design, incident review rigor, and postmortem-driven improvements
  • Provide architectural guidance and mentorship to senior engineers and technical leaders across teams
  • Balance technical strategy, hands-on engineering judgment, and cross-functional influence to drive measurable improvements in site stability
  • Help shape how LinkedIn builds and operates resilient systems as the platform continues to scale
  • Drive the strategy for applying LLMs to alert triage, root cause analysis, and incident summarization at scale, ensuring systems are explainable, auditable, and safe to operate autonomously in Ring0/Ring1 environments.

Benefits

  • annual performance bonus
  • stock
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service