Staff Engineer, Hardware Reliability

LinkedInSunnyvale, CA
Hybrid

About The Position

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. This role will be based in Sunnyvale, CA. We are looking for a highly skilled, self-motivated Staff Engineer to join our Hardware Capacity Engineering (HCE) team and help us scale and sustain the infrastructure that powers LinkedIn. HCE qualifies, integrates, and operates the full range of hardware in our on-prem data centers , such as general-purpose compute, GPU/accelerator, storage, and networking platforms, across a large-scale, multi-vendor, multi-generation fleet. This role spans both bringing new platforms into production and keeping our existing fleet healthy, performant, and reliable, backed by the software, firmware automation, and fleet-health systems the team builds. In this role, you will identify requirements and the best-suited hardware platform or solution, integrate that solution into our on-prem data center environment, and help operate and continuously improve the existing fleet at scale. You will build software and automation that make the fleet observable, performant, and reliable, and partner closely with SRE, software engineering, AI/ML, and hardware vendors.

Requirements

  • BS in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • 6+ years of experience working in Linux-based infrastructure, systems, or hardware engineering.
  • 4+ years of experience with hardware troubleshooting, systems engineering, and performance analysis.
  • Experience developing software or automation (AI-assisted or otherwise) for infrastructure at scale.

Nice To Haves

  • Experience with x86 server architecture and multi-vendor hardware, BMC/BIOS and firmware (IPMI/Redfish).
  • Experience qualifying and integrating new hardware platforms and operating them across a large-scale, diverse, multi-generation fleet.
  • Experience with hardware provisioning, imaging/OS, and lifecycle or inventory systems
  • Experience building fleet health, observability, or reliability tooling (fault detection, SMART and telemetry analysis, data-driven thresholds).
  • Experience benchmarking with common tools such as SPEC, SPECpower, fio, unixbench, and similar.
  • Experience with storage devices and performance engineering (NVMe/SSD/HDD) and/or distributed/parallel filesystems (GPFS, HDFS).
  • Experience with GPU/accelerator platforms and the ML training stack (NCCL/collective communications, CUDA) and high-performance networking (InfiniBand/RDMA, RoCE).
  • Experience with Kubernetes and containerized workloads.
  • Experience working with hardware vendors on both designing a solution and troubleshooting issues.
  • Demonstrated experience putting together summary reports and visual presentations of the results of benchmarks and performance tests for technical and executive audiences.
  • Familiarity with HPC/Machine Learning environments and solutions.

Responsibilities

  • Collaborate with LinkedIn engineering teams to collect requirements for LinkedIn applications and provide guidance on selecting the best-suited hardware platforms and solutions across general-purpose compute, GPU/accelerator, storage, and networking.
  • Design test environments and testing scenarios; benchmark compute, storage, and power (e.g., SPEC, SPECpower, etc) and provide detailed analysis of qualification and performance results for varied audiences, including engineers and senior leadership.
  • Qualify and integrate new server platforms and components end-to-end, working with hardware vendors on optimal configurations and driving the full integration process.
  • Work jointly with other teams on cost and TCO analysis for proposed solutions, present them to decision makers, and define SLAs and technical standards with partner teams.
  • Qualify BIOS, BMC, and component firmware.
  • Drive fleet-wide upgrade programs.
  • Support and improve the reliability of our existing large-scale, diverse fleet including fault detection and remediation, firmware management, and OS and security compliance.
  • Design, build (AI-assisted) and own automation for hardware qualification, provisioning, lifecycle, and fleet health monitoring and telemetry analysis.
  • Contribute to AI/ML infrastructure performance and reliability of GPU platforms, InfiniBand/RDMA as part of the team’s broader scope.
  • Troubleshoot complex hardware, firmware, kernel, and platform issues across the fleet, and lead critical production incident response.

Benefits

  • Generous health and wellness programs
  • Time away for employees of all levels
  • Annual performance bonus
  • Stock
  • Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service