Software Development Engineer(Distributed Systems)

Workday•Pleasanton, CA
•$123,900 - $222,000•Hybrid

About The Position

The Data Platform and Observability Engineering (DPOE) team is building Workday’s next-generation, multi-petabyte-scale Observability Platform. We are a distributed team responsible for the libraries, distributed services, and infrastructure that power observability across Workday. Our stack includes Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, Elasticsearch, and LangSmith for LLM/agentic tracing and evaluation. Together, we support traces, metrics, logs, and LLM/agent observability across Workday workloads at scale. As a Software Development Engineer, you will build and scale core features of Workday’s distributed tracing platform. Working with ClickHouse, Grafana Tempo, and a modern big-data pipeline (Kafka, Spark/Flink, Iceberg, S3) on AWS, you will deliver high-quality, high-performance code capable of handling massive data loads.You will optimize low-latency infrastructure and help shape the future of Observability AI by extending distributed tracing to LLM and agentic workflows using LangSmith and LangChain, enabling visibility into agent execution, tool calls, and prompt/response chains. This is a hands-on role for an engineer who excels at building resilient backend services and is eager to grow their expertise in distributed systems and big data while contributing to the future of Observability AI.

Requirements

  • 5+ years experience in software development engineering.
  • 3+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.
  • 5+ years experience with at least two of the following programming languages Java, Python, Go, including experience in writing production-level code for distributed systems.
  • Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.

Nice To Haves

  • Strong ability in Algorithmic to build highly efficient and scalable solutions for complex high-throughput data ingestion and sub-second query performance challenges.
  • Solid experience in API Development, including an understanding of gRPC, REST, and OpenTelemetry (OTLP), with practical experience designing and building scalable distributed APIs for observability data.
  • Strong understanding of Code Testing methodologies, such as distributed load testing and integration testing, and experience contributing to end-to-end telemetry pipeline testing and CI/CD automation.
  • Solid understanding of Distributed Systems Software principles, including data partitioning, eventual consistency, and fault tolerance mechanisms, with hands-on experience in Kafka, Spark, Flink, or ClickHouse.
  • Experience implementing and maintaining High Availability strategies for critical distributed systems, including multi-AZ deployments, robust retry mechanisms, and automated failover.
  • Practical experience with Large Scale Data Processing technologies and frameworks such as Apache Kafka, Spark, Flink, and Apache Iceberg within complex distributed architectures.
  • Good understanding of Large Scale Systems design principles, including distributed data sharding, replication, and query optimization, and experience working on observability pipelines or data lake platforms.
  • Strong understanding of Object-Oriented Design (OOD) principles and architectural patterns for building highly scalable and maintainable distributed systems.
  • Experience with Source Control Management (SCM) tools such as Git and GitHub/Bitbucket, and following best practices for collaborative distributed development workflows.
  • Strong understanding of System Security principles and best practices relevant to securing distributed environments, including mutual TLS (mTLS), multi-tenant authorization (authz), and data encryption.
  • Proven ability to actively collaborate within and across distributed software development teams and contribute constructively to architectural discussions and system designs.
  • Strong skills in creating Technical Writing Documentation for runbooks, system design specs, and API documentation related to distributed systems architecture and design.

Responsibilities

  • Build and scale core features of Workday’s distributed tracing platform.
  • Deliver high-quality, high-performance code capable of handling massive data loads.
  • Optimize low-latency infrastructure.
  • Extend distributed tracing to LLM and agentic workflows using LangSmith and LangChain.
  • Enable visibility into agent execution, tool calls, and prompt/response chains.

Benefits

  • Workday Bonus Plan
  • role-specific commission/bonus
  • annual refresh stock grants
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service