Sr. Product Manager - Observability

WorkdayPleasanton, CA
$140,600 - $252,000Hybrid

About The Position

We are looking for a hands-on, technical Senior Product Manager (P4) to own and drive our Observability strategy, with a strong emphasis on Distributed Tracing. You will define the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on building AI-powered detection and triage capabilities that reduce time-to-detect and time-to-resolve production issues. This role requires someone comfortable diving deep into technical architecture discussions, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products.

Requirements

  • 8+ years of product management experience, with meaningful time spent in Observability, Monitoring, APM, or related infrastructure/developer tooling domains
  • Deep, demonstrable expertise in Distributed Tracing concepts and technologies (e.g., OpenTelemetry, Jaeger, Zipkin, trace sampling, span/context propagation)
  • Strong technical background — comfortable reading code, understanding system architecture, APIs, and engaging directly in technical design discussions
  • Concrete, hands-on understanding of AI/ML techniques applied to detection and triage, including: Anomaly detection methods (statistical thresholds, time-series forecasting models, change-point/seasonality detection), Alert correlation and clustering techniques (embedding/vector similarity, graph-based dependency analysis, unsupervised clustering), LLM-based summarization and reasoning for incident root-cause analysis and runbook suggestion, Model evaluation frameworks (precision/recall, false-positive/false-negative trade-offs, drift monitoring)
  • Experience partnering directly with ML engineering teams to translate detection and triage requirements into shippable models and features
  • Proven track record of shipping complex, technical products from concept to launch
  • Excellent written and verbal communication skills; ability to translate complex technical and ML concepts for varied audiences
  • Strong analytical and data-driven decision-making skills

Nice To Haves

  • Prior experience as a software engineer, SRE, ML engineer, or data scientist before transitioning to product management
  • Familiarity with observability platforms (Datadog, New Relic, Honeycomb, Grafana, Splunk, Dynatrace, etc.) and their AI/ML-driven detection features
  • Direct experience productionizing ML models (e.g., anomaly detectors, classifiers, LLM-based tools) for operational/incident-management use cases
  • Familiarity with vector databases/embedding search, graph-based correlation techniques, or LLM prompt/agent design for triage automation
  • Background in large-scale, cloud-native, microservices-based systems

Responsibilities

  • Own the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.)
  • Define and drive the roadmap for AI-enabled anomaly detection, including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs)
  • Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement
  • Define requirements for LLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns
  • Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on-call engineers can trust and act on AI-generated recommendations
  • Partner closely with engineering teams to define technical requirements, evaluate architectural trade-offs, and make hands-on contributions to product decisions (e.g., reviewing design docs, understanding OpenTelemetry/OTel standards, tracing protocols, and instrumentation approaches)
  • Define and track success metrics (detection precision/recall, MTTD, MTTR, alert noise reduction, triage automation rate, trace coverage) to measure product impact
  • Conduct customer and internal stakeholder research to identify pain points in debugging, alert fatigue, and incident triage
  • Write clear, detailed product requirements, user stories, and specs; work closely with design, engineering, and data science to bring them to life
  • Stay current on industry trends in observability (OpenTelemetry, eBPF-based tracing, service mesh telemetry) and applied AI/ML for anomaly detection, incident correlation, and LLM-based operational tooling

Benefits

  • Workday Bonus Plan or a role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service