Principal, Software Engineer (Distributed Systems)

WorkdayPleasanton, CA
$187,100 - $334,300Hybrid

About The Position

Workday is seeking a Principal Software Engineer specializing in Distributed Systems to join their Data Platform and Observability team. This role involves architecting and developing large-scale distributed data systems that support critical Workday products and provide real-time insights across Workday’s platforms, infrastructure, and applications. The team handles hundreds of terabytes of data for core Workday products, including HCM, Fins, AI/ML, internal data products, and Observability. The ideal candidate will enjoy writing efficient software and tuning/scaling large distributed systems, tackling challenges at massive scale for over 10,000 global customers, and working with world-class engineers to develop next-generation distributed systems platforms. The Messaging, Streaming, and Caching team, specifically, is a full-service Distributed Systems Engineering team responsible for architecting and providing async messaging, streaming, and NoSQL platforms and solutions that power Workday products. They develop client libraries, SDKs, and automation for deploying and operating hundreds of clusters, ensuring resiliency and operational excellence. Team members play a key role in improving services and encouraging their adoption within Workday's infrastructure, both in private and public clouds, designing and building new capabilities from inception to deployment.

Requirements

  • 12+ years experience in software development engineering.
  • 6+ years experience specifically focused on designing, building, and operating distributed systems like Redis, Kafka, RabbitMQ or NoSQL solutions.
  • 5+ experience in designing and implementing complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.99% uptime) and fault tolerance.
  • 8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go, C/C++), including experience in writing production-level code for distributed systems.
  • Expertise with configuration management using Chef and service deployment on Kubernetes via Helm and ArgoCD
  • Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.

Nice To Haves

  • Expert-level ability in Algorithmic Thinking, including CAP theorem, queuing theory, consensus protocols etc, to architect highly efficient and scalable solutions for complex distributed systems implementations.
  • Deep expertise in API and Client Library Development, including understanding of RESP protocol, Kafka wire protocol etc and extensive experience in designing and building API layer as well as client libraries.
  • Experience building cloud native controllers for distributed systems. Familiarity with operator like Strimzi would be a bonus
  • Strong understanding of modern Code Testing methodologies like consistency / linearizability testing, and experience in leading chaos and and fault injection testing strategies.
  • Deep understanding of Distributed Systems Software principles, including fault tolerance, high availability, and extensive experience in replication / sharding techniques
  • Proven ability to design and implement High Availability strategies for critical distributed systems, including global replication with 99.99% SLOs.
  • Extensive experience with Large Scale Data Processing technologies and frameworks such as Kafka/Redis/RabbitMQ/Spark/Flink etc within complex distributed architectures.
  • Strong understanding of Object-Oriented Design (OOD) principles and architectural patterns for building highly scalable and maintainable distributed systems.
  • Extensive experience with SCM and CI/CD tools such as Git, Jenkins, Harness etc and establishing best practices for collaborative distributed development workflows.
  • Strong understanding of System Security principles and best practices relevant to securing complex distributed environments, including implementing authentication and authorization modules.
  • Experience learning complex open source service internals via code inspection.
  • Proven ability to lead Team Collaboration within and across distributed software development teams and drive architectural direction.
  • Strong skills in creating Technical Writing Documentation and presenting to senior leadership as well as architects within the company.

Responsibilities

  • Design, build, and enhance critical distributed services, including Kafka, Redis, RabbitMQ etc.
  • Design, develop, build, deploy and maintain core distributed services using a combination of open source and proprietary stacks across diverse infrastructure environments (Kubernetes, OpenStack, Bare Metal, etc.)
  • Design and develop core software modules for streaming, messaging and caching.
  • Build observability modules, alerts and automation for Dashboard lifecycle management for the distributed services.
  • Build, deploy and operate infrastructure components in production environments.
  • Champion all aspects of streaming, messaging and caching with a focus on resiliency and operational excellence.
  • Evaluate and implement new open-source and cloud-native tools and technologies as needed.
  • Participate in the on-call rotation to support the distributed systems platforms.
  • Manage and optimize Workday distributed services in AWS, GCP & Private cloud env.

Benefits

  • Workday Bonus Plan
  • Role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service