Principal Distributed Systems Engineer - Observability

WorkdayPleasanton, CA
$187,100 - $334,300Hybrid

About The Position

Workday is seeking a Principal Distributed Systems Engineer specializing in Observability. This role is crucial for developing Workday's next-generation, multi-petabyte scale Observability Platform. The team owns the libraries, distributed services, and infrastructure for ingestion, storage, and query across the observability stack, including Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch. The work directly influences how the company detects, diagnoses, and predicts operational issues at scale. The position involves owning the technical vision and architecture for distributed tracing, built on ClickHouse and/or Grafana Tempo, supported by a big-data pipeline on AWS. It's a hands-on, high-autonomy role for an engineer to design and build multi-petabyte, low-latency tracing infrastructure end-to-end. The role also involves defining the future of Observability AI, utilizing traces, logs, and metrics for automated root-cause analysis, anomaly detection, and AI-driven incident triage. The engineer will set technical direction, mentor senior and staff engineers, and serve as the primary architect and escalation point for the tracing subsystem.

Requirements

  • 14+ years experience in software development engineering.
  • 6+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.
  • 8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go), including experience in writing production-level code for distributed systems.
  • Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.
  • Expert-level ability in Algorithmic Thinking, including [insert specific advanced algorithms or data structures relevant to distributed systems], to architect highly efficient and scalable solutions for complex
  • Deep expertise in API Development, including understanding of advanced API protocols or architectural patterns
  • Deep understanding of Distributed Systems Software principles, like distributed consensus or fault tolerance mechanisms
  • Proven ability to design and implement High Availability strategies for critical distributed systems
  • Extensive experience with Large Scale Data Processing technologies and frameworks
  • Deep understanding of Large Scale Systems design principles like distributed data management or scalability strategies
  • Strong understanding of System Security principles and best practices relevant to securing complex distributed environments
  • Proven ability to lead Team Collaboration within and across distributed software development teams and drive architectural direction
  • Strong skills in creating Technical Writing Documentation and Presentation

Nice To Haves

  • Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.

Responsibilities

  • Architect and build Workday's distributed tracing platform on ClickHouse/Tempo, designed for multi-petabyte scale ingestion and sub-second interactive query performance.
  • Own the big-data pipeline feeding tracing data — Kafka-based ingestion, Spark/Flink stream and batch processing, and Iceberg-on-S3 storage — including schema design, partitioning, compaction, and lifecycle management.
  • Drive performance and scaling across ingestion and query paths: storage format optimization (Parquet/Iceberg), compression strategy, partitioning/indexing, and query engine tuning under real production load.
  • Lead HA/DR design for tracing services — multi-region/multi-AZ resilience, failover, backup/restore, and recovery time/point objectives appropriate to a tier-1 platform.
  • Design security architecture for the platform, including authentication/authorization (authn/authz) for multi-tenant data access across ingestion and query layers.
  • Own operational excellence for distributed tracing: monitoring, logging, alerting, capacity planning, and participation in an on-call rotation for the platform.
  • Evaluate and introduce new technologies — open source and cloud-native — that materially improve the platform's scalability, cost efficiency, or capability.
  • Shape the future of Observability AI: partner with ML/AI stakeholders to define how tracing data feeds automated anomaly detection, root-cause analysis, and AI-assisted incident management.
  • Evangelize the platform: publish best practices, mentor engineers across DPOE and partner teams, and act as a technical thought leader for the modern observability/data stack internally.
  • Operate with high autonomy in a fast-moving, ambiguous environment — setting technical direction with minimal oversight while aligning with broader platform strategy.

Benefits

  • Workday Bonus Plan
  • Role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service