Sr. Principal Architect - Observability

Workday•Pleasanton, CA
•$221,600 - $392,600•Hybrid

About The Position

Workday is seeking a Sr. Principal Architect for Observability. This is a hands-on role responsible for the architecture, standards, and best practices for all things Observability at Workday. The individual will work closely with Observability Platform engineers, product managers, and stakeholders to evolve the platform architecture, contribute to the Observability platform, and act as an evangelist. The role involves providing architectural leadership for Workday’s Observability Platform and Services, acting as the domain expert across all pillars of Observability. Key responsibilities include driving agentic experiences for incident detection, triage, and recovery by owning and driving the Observability MCP and Observability Agents, ensuring a seamless Observability experience across data domains (logs, metrics, traces, alerts) and different infrastructure types, driving standardization of Observability tooling, developing and publishing standards, and leveraging open-source specifications like OpenTelemetry. The role also requires close collaboration with Infrastructure, Technology Operations, and Product teams, and promoting and evangelizing Observability to all Workday engineering teams, establishing standard processes, frameworks, and assisting with onboarding engineering teams.

Requirements

  • BS/MS in Computer Science or a related technical field
  • At least 10+ years of software engineering/architect background and a proven track record of delivering high quality products at massive scale
  • 10+ years' experience writing high quality code (Java/Scala/Go/Python) and be able to design systems that meet the business needs.
  • Hands-on expertise in Observability platforms like Prometheus, Mimir, Elasticsearch, Clickhouse, VictoriaMetrics, Jaeger/Zipkin etc.

Nice To Haves

  • Extensive experience with data collection agents, loggers, and visualizing and correlating data sets from multiple domains.
  • Expertise in LLM orchestration framework (LangChain etc,), Observability for AI Agents (Latency, quality, online/offline evals, Track multi-turn conversations, tool/agent invocations etc)
  • Experience building Observability agents that help with Triage, Root Cause Analysis using open standards like Skills, MCPs and popular orchestration frameworks like LangChain etc.
  • Deployed Observability across the OSI stack.
  • Expertise in low-latency/high throughput data processing capabilities a big plus.
  • Open Source contributions to popular O11y frameworks & platforms are a big plus.
  • Outstanding presentation skills to both technical and executive audiences.
  • Strong communication skills both written and verbal.

Responsibilities

  • Provide architectural leadership for Workday’s Observability Platform and Services.
  • Act as the domain expert across all the pillars of Observability.
  • Drive Agentic experiences for incident detection, triage, and recovery by owning & driving the Observability MCP and Observability Agents.
  • Ensure a seamless Observability experience across data domains (logs, metrics, traces, alerts) and different infrastructure types (Virtualized, Kubernetes, Bare metal, etc.).
  • Drive standardization of the Observability tooling, develop and publish standards, leverage open source specifications (OpenTelemetry etc).
  • Closely collaborate with the Infrastructure, Technology Operations & Product teams.
  • Promote & Evangelize Observability to all Workday engineering teams.
  • Establish standard processes, frameworks, and help onboard engineering teams.

Benefits

  • Workday Bonus Plan or a role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service