At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. This role may be remote or hybrid. At LinkedIn, hybrid roles are performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Remote roles are performed from the designated home work location upon time of hire, and any changes to this home work location requires a review of remote status and approval. LinkedIn’s Reliability Infrastructure team is responsible for defining and driving the reliability strategy, standards, and practices that keep LinkedIn’s most critical systems stable, resilient, and available at massive scale. As a Principal Staff Software Engineer, Reliability Infrastructure, you will serve as a senior technical authority for reliability across LinkedIn Engineering. You will help define how critical services are designed, built, operated, and measured, partnering broadly across infrastructure and product engineering teams to improve resiliency, reduce incidents, and raise the reliability bar across the company. A key focus of this role is driving the adoption and evolution of LinkedIn’s service criticality framework, including reliability expectations for the most business-critical systems. You will help classify services based on criticality and blast radius, define appropriate reliability standards, and influence system architecture to ensure the right levels of availability, redundancy, observability, and failure handling are in place. As AI-assisted software development, agent-based automation, and autonomous operational systems become more prevalent, this role will also help define how LinkedIn safely builds and operates reliable AI-enabled systems. You will shape standards for evaluating, deploying, monitoring, and governing AI-generated code and agentic workflows, ensuring that automation introduced into critical environments is observable, explainable, auditable, and designed with appropriate safeguards, rollback mechanisms, and human oversight. This is not a traditional SRE role focused on operating a single service or team. It is a company-wide technical leadership role for someone with deep distributed systems expertise, strong reliability judgment, and the ability to influence architecture and engineering practices across large organizations.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal
Education Level
Associate degree