Sr. Software Engineer - Distributed Systems

Workday•Boulder, CA
•Hybrid

About The Position

Workday is seeking a Sr. Software Engineer for their Data Platform and Observability team. This role involves designing, building, and improving critical observability services like Monitoring, Logging, Alerting, and Tracing. The team focuses on a large-scale distributed data and observability platform that ensures Workday's services, including AI agents and automated workflows, are fast, reliable, and trustworthy in production. They handle massive amounts of data daily and are integrating AI-driven approaches for anomaly detection, intelligent alerting, and automated root-cause analysis. The role offers the opportunity to solve complex challenges at massive scale across private and public cloud environments, contributing to the next generation of AI-driven products.

Requirements

  • 8 years of software engineering/design experience.
  • 7 years of coding in Python, Go, or Java.
  • Experience with Linux.
  • BS in Computer Science or a related technical field, or equivalent experience.

Nice To Haves

  • Ability to design, maintain, and optimize Time-Series DBs, document stores, and log collection/management systems.
  • Solid understanding of distributed tracing concepts (context propagation, spans, sampling) and hands-on experience with tracing tooling.
  • Public cloud experience (AWS/GCP), including native observability tooling such as AWS CloudWatch and GCP Cloud Operations (Stackdriver).
  • Experience with containerization and infrastructure automation (Docker, Kubernetes, Ansible, Chef, Terraform).
  • Experience with service mesh, Prometheus, and cloud-native technologies.
  • Hands-on experience with LLM orchestration frameworks (e.g., LangChain, LlamaIndex, Semantic Kernel) and a working knowledge of agentic AI build patterns — how agents plan, call tools, hold state, and hand off work across multi-step chains.
  • A knack for spotting where AI systems quietly go wrong in production: a model call that's suddenly slower than usual, a token bill creeping up for no clear reason, or answers that subtly drift in quality over time. You know how to build the dashboards, alerts, and instrumentation that catch these early — turning "the AI feels off" into a measurable, debuggable signal.
  • Excellent interpersonal, technical, and communication skills.
  • Ability to prioritize multiple tasks in a fast-paced environment.
  • MS Degree a plus.

Responsibilities

  • Design, build, and improve critical observability services: Monitoring, Logging, Alerting, and Tracing.
  • Own distributed tracing end to end — instrumenting services, propagating context across service boundaries, and using trace data to understand system behavior and diagnose issues in complex, multi-hop request paths.
  • Build data capture and collection services using the latest technologies across multiple infrastructure types (Kubernetes, Docker, OpenStack, bare metal, etc.).
  • Design and develop core software modules for real-time and batch data processing.
  • Build metrics ingestion pipelines, alert definitions, and automation for dashboard lifecycle management.
  • Instrument AI-powered and automated workflows for performance, cost, and quality, including multi-step processes and tool/service orchestration.
  • Explore and apply AI-driven observability techniques — anomaly detection, intelligent alerting, and automated root-cause analysis — to reduce noise and speed up incident response.
  • Partner with product and application teams to define SLOs/SLIs and production-readiness criteria for new services and automated workflows.
  • Build infrastructure components and deploy them in production.
  • Work across all aspects of observability with a keen eye for data quality, data integrity, and data availability.
  • Evaluate and implement new open-source and cloud-native tools and technologies as needed.
  • Participate in the on-call rotation supporting the observability platform.

Benefits

  • Workday Bonus Plan or a role-specific commission/bonus
  • Annual refresh stock grants
  • Comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service