Workday is seeking a Principal Distributed Systems Engineer specializing in Observability. This role is crucial for developing Workday's next-generation, multi-petabyte scale Observability Platform. The team owns the libraries, distributed services, and infrastructure for ingestion, storage, and query across the observability stack, including Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch. The work directly influences how the company detects, diagnoses, and predicts operational issues at scale. The position involves owning the technical vision and architecture for distributed tracing, built on ClickHouse and/or Grafana Tempo, supported by a big-data pipeline on AWS. It's a hands-on, high-autonomy role for an engineer to design and build multi-petabyte, low-latency tracing infrastructure end-to-end. The role also involves defining the future of Observability AI, utilizing traces, logs, and metrics for automated root-cause analysis, anomaly detection, and AI-driven incident triage. The engineer will set technical direction, mentor senior and staff engineers, and serve as the primary architect and escalation point for the tracing subsystem.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal