Sr. Product Manager - Observability

WorkdayPleasanton, CA
Hybrid

About The Position

Workday is seeking a hands-on, technical Senior Product Manager (P4) to own and drive the Observability strategy, with a strong emphasis on Distributed Tracing. This role involves defining the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on building AI-powered detection and triage capabilities to reduce time-to-detect and time-to-resolve production issues. The ideal candidate is comfortable diving deep into technical architecture, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products. The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform, owning the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack. This platform is critical for engineering, SRE, and product teams at Workday to monitor system performance and identify issues before they impact customers. The roadmap directly influences how the company detects, diagnoses, and predicts operational issues at scale, transforming observability into a proactive, AI-assisted safety net.

Requirements

  • 8+ years of product management experience, with meaningful time spent in Observability, Monitoring, APM, or related infrastructure/developer tooling domains
  • Deep, demonstrable expertise in Distributed Tracing concepts and technologies (e.g., OpenTelemetry, Jaeger, Zipkin, trace sampling, span/context propagation)
  • Strong technical background — comfortable reading code, understanding system architecture, APIs, and engaging directly in technical design discussions
  • Concrete, hands-on understanding of AI/ML techniques applied to detection and triage , including: Anomaly detection methods (statistical thresholds, time-series forecasting models, change-point/seasonality detection), Alert correlation and clustering techniques (embedding/vector similarity, graph-based dependency analysis, unsupervised clustering), LLM-based summarization and reasoning for incident root-cause analysis and runbook suggestion, Model evaluation frameworks (precision/recall, false-positive/false-negative trade-offs, drift monitoring)
  • Experience partnering directly with ML engineering teams to translate detection and triage requirements into shippable models and features
  • Proven track record of shipping complex, technical products from concept to launch
  • Excellent written and verbal communication skills; ability to translate complex technical and ML concepts for varied audiences
  • Strong analytical and data-driven decision-making skills

Nice To Haves

  • Prior experience as a software engineer, SRE, ML engineer, or data scientist before transitioning to product management
  • Familiarity with observability platforms (Datadog, New Relic, Honeycomb, Grafana, Splunk, Dynatrace, etc.) and their AI/ML-driven detection features
  • Direct experience productionizing ML models (e.g., anomaly detectors, classifiers, LLM-based tools) for operational/incident-management use cases
  • Familiarity with vector databases/embedding search, graph-based correlation techniques, or LLM prompt/agent design for triage automation
  • Background in large-scale, cloud-native, microservices-based systems

Responsibilities

  • Own the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.)
  • Define and drive the roadmap for AI-enabled anomaly detection, including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs)
  • Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement
  • Define requirements for LLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns
  • Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on-call engineers can trust and act on AI-generated recommendations
  • Partner closely with engineering teams to define technical requirements, evaluate architectural trade-offs, and make hands-on contributions to product decisions (e.g., reviewing design docs, understanding OpenTelemetry/OTel standards, tracing protocols, and instrumentation approaches)
  • Define and track success metrics (detection precision/recall, MTTD, MTTR, alert noise reduction, triage automation rate, trace coverage) to measure product impact
  • Conduct customer and internal stakeholder research to identify pain points in debugging, alert fatigue, and incident triage
  • Write clear, detailed product requirements, user stories, and specs; work closely with design, engineering, and data science to bring them to life
  • Stay current on industry trends in observability (OpenTelemetry, eBPF-based tracing, service mesh telemetry) and applied AI/ML for anomaly detection, incident correlation, and LLM-based operational tooling

Benefits

  • Workday Bonus Plan or a role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service