Senior Staff Software Engineer, AI Infrastructure

LinkedInSunnyvale, CA
Hybrid

About The Position

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Join us to build the platforms that enable LinkedIn to evaluate, monitor, and continuously improve machine learning models at scale. Our AI systems power recommendations, search, ads, LLMs, computer vision, and other intelligent experiences used across LinkedIn. The Model Evaluation team develops robust, scalable frameworks that empower engineers and researchers to rigorously quantify model quality, conduct comparative analysis against established baselines, proactively identify performance regressions, and seamlessly bridge the gap between offline evaluation metrics and real-world production outcomes. The Model Observability team engineers robust, highly scalable infrastructure that delivers continuous, real-time insights into model performance and behavior in production. We empower teams to proactively detect and diagnose critical issues—including model drift, training-serving skew, degradation in data quality, and shifts in score distributions—ensuring that our AI systems remain reliable, trustworthy, and performant at scale. As a Sr. Staff Software Engineer, you will help define and build LinkedIn’s next generation of Model Evaluation and Observability infrastructure, solving complex distributed systems and ML platform problems while influencing how AI systems are evaluated and understood across the company.

Requirements

  • BS/BA in Computer Science or related technical field or equivalent technical experience
  • 5+ years of industry experience in software design, development, and algorithm-related solutions
  • 5+ years of experience programming in languages such as Python, C++, Java, Go, Rust, or Scala
  • 2+ years of experience as an architect, technical lead, or in another technical leadership position
  • 5+ years of experience building large-scale infrastructure, machine learning systems, or distributed systems
  • Hands-on experience designing and developing distributed systems or other large-scale production platforms

Nice To Haves

  • MS or PhD in Computer Science or related technical discipline
  • 10+ years of experience in software design and development, including significant experience in technical leadership positions
  • 5+ years of experience designing and building large-scale distributed systems and production infrastructure.
  • Experience building machine learning infrastructure, model lifecycle platforms, or large-scale production ML systems.
  • Experience building model evaluation, model monitoring, ML observability, experimentation, model validation, or model quality infrastructure.
  • Experience with generative recommendation architectures, including LLM/SLM-based rankers, semantic ID representations, and evaluation of sequence-to-sequence or autoregressive ranking models.
  • Experience designing platforms that collect and process model outputs, metrics, metadata, telemetry (OTEL or OpenInferenceTelemetry), or other production ML signals at scale.

Responsibilities

  • Own the technical strategy and architecture for large-scale Model Evaluation and Observability infrastructure, developing solutions that span multiple product lines and AI use cases.
  • Design highly available, distributed architectures to ingest, process, and analyze high-volume telemetry data from a variety of models, encompassing recommendation and ranking, machine learning, LLMs and generative AI systems.
  • Build scalable model evaluation platforms that enable ML engineers and researchers to measure model quality, compare models, identify regressions, and understand model behavior across experimentation and production environments.
  • Lead the diagnosis and resolution of complex, cross-team performance bottlenecks, data quality issues, and systemic reliability challenges in the ML lifecycle.
  • Define and implement "observability-by-default" frameworks that enable ML engineers to iterate faster by seamlessly bridging the gap between experimentation, offline evaluation, and production reliability.
  • Build capabilities for identifying and diagnosing issues such as model regressions, drift, training-serving skew, score-distribution changes, and data-quality problems.
  • Improve developer productivity by making it easier for teams to evaluate, monitor, and diagnose production ML systems.
  • Mentor and influence engineers across the organization, establish strong engineering practices, and raise the technical bar for large-scale ML infrastructure.
  • Serve as a technical leader across multiple Model Evaluation and Model Observability initiatives, driving architecture and execution across organizational boundaries.
  • Anticipate future scale and complexity requirements, proactively evolving our architecture to handle increasing data volumes, diverse model types, and evolving compliance/governance standards.

Benefits

  • annual performance bonus
  • stock
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service