Senior Software Engineer(Distributed Systems)

WorkdayPleasanton, CA
$160,100 - $285,100Hybrid

About The Position

As a Senior Software Development Engineer, you will own the technical design and execution for core components of Workday’s distributed tracing platform, built on ClickHouse and/or Grafana Tempo and backed by a big-data pipeline (Kafka, Spark/Flink, Iceberg, S3). This is a hands-on role where you will tackle multi-petabyte scale challenges, optimize low-latency infrastructure, and help lay the technical groundwork for Observability AI. You will drive engineering excellence within the team, mentor other developers, and act as a domain expert for tracing ingestion, storage, and query performance.

Requirements

  • 8+ years experience in software development engineering.
  • 4+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.
  • 5+ years experience with at least two of the following programming languages Java, Go, Scala, Python, including experience in writing production-level code for distributed systems.
  • Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.
  • Strong ability in Algorithmic Thinking, including to build highly efficient and scalable solutions for complex high-throughput data ingestion and sub-second query performance challenges.
  • Solid experience in API Development, including an understanding of gRPC, REST, and OpenTelemetry (OTLP), with practical experience designing and building scalable distributed APIs for observability data.
  • Strong understanding of Code Testing methodologies, such as distributed load testing and integration testing, and experience contributing to end-to-end telemetry pipeline testing and CI/CD automation.
  • Solid understanding of Distributed Systems Software principles, including data partitioning, eventual consistency, and fault tolerance mechanisms, with hands-on experience in Kafka, Spark, Flink, or ClickHouse.
  • Experience implementing and maintaining High Availability strategies for critical distributed systems, including multi-AZ deployments, robust retry mechanisms, and automated failover.
  • Practical experience with Large Scale Data Processing technologies and frameworks such as Apache Kafka, Spark, Flink, and Apache Iceberg within complex distributed architectures.
  • Good understanding of Large Scale Systems design principles, including distributed data sharding, replication, and query optimization, and experience working on observability pipelines or data lake platforms.
  • Strong understanding of Object-Oriented Design (OOD) principles and architectural patterns for building highly scalable and maintainable distributed systems.
  • Experience with Source Control Management (SCM) tools such as Git and GitHub/Bitbucket, and following best practices for collaborative distributed development workflows.
  • Strong understanding of System Security principles and best practices relevant to securing distributed environments, including mutual TLS (mTLS), multi-tenant authorization (authz), and data encryption.
  • Proven ability to actively collaborate within and across distributed software development teams and contribute constructively to architectural discussions and system designs.
  • Strong skills in creating Technical Writing Documentation for runbooks, system design specs, and API documentation related to distributed systems architecture and design.

Nice To Haves

  • Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred

Responsibilities

  • Own the technical design and execution for core components of Workday’s distributed tracing platform.
  • Tackle multi-petabyte scale challenges.
  • Optimize low-latency infrastructure.
  • Help lay the technical groundwork for Observability AI.
  • Drive engineering excellence within the team.
  • Mentor other developers.
  • Act as a domain expert for tracing ingestion, storage, and query performance.

Benefits

  • Workday Bonus Plan
  • role-specific commission/bonus
  • annual refresh stock grants
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service