Software Engineer, Distributed Systems

eBay•San Jose, CA
•$172,000 - $229,600•Remote

About The Position

The Observability Platform team builds and operates the infrastructure that helps eBay teams monitor, fix, and improve the reliability of large-scale distributed systems. This platform supports the telemetry and reliability needs of thousands of microservices across eBay and operates at hyperscale, processing billions of time series and petabytes of log data using modern open-source technologies including Prometheus, ClickHouse, OpenTelemetry, and related tools. As a Software Engineer on this team, you will design and build scalable distributed systems that power metrics, logs, traces, and related observability workflows across the stack—from ingestion and storage to query and visualization. You will partner closely with SREs, platform engineers, and service owners to solve complex reliability challenges, improve operational excellence, and help shape the future of observability at eBay. This role offers the opportunity to work on critical systems, contribute to open-source technologies, and grow through direct exposure to some of eBay’s most complex infrastructure challenges. This role also includes participation in the team’s on-call rotation in support of production reliability.

Requirements

  • 7+ years of experience in software engineering, infrastructure engineering, or a closely related field
  • Strong programming skills in Golang or another systems-level language, with experience building reliable backend or infrastructure services
  • Deep understanding of distributed systems concepts such as fault tolerance, scalability, and system reliability
  • Hands-on experience deploying and operating containerized services in Kubernetes or similar cloud-native environments
  • Solid understanding of observability domains including metrics, logs, and traces

Nice To Haves

  • Experience with tools such as Prometheus, Grafana, OpenTelemetry, ClickHouse, or similar technologies
  • Familiarity with time-series systems, high-throughput data pipelines, open-source infrastructure, or React/JavaScript is a plus

Responsibilities

  • Design and deliver scalable, fault-tolerant observability infrastructure that improves reliability while reducing operational overhead for platform and engineering teams
  • Build and optimize high-throughput services for ingesting, transforming, storing, and querying telemetry data across logs, metrics, and traces
  • Strengthen the resilience of Kubernetes-based production systems through self-healing, autoscaling, and robust operational design
  • Partner with SREs, platform teams, and service owners to translate observability needs into tools and platform capabilities that improve incident response and operational excellence
  • Contribute to architecture reviews, production readiness discussions, and post-incident findings to drive continuous improvement across the platform
  • Expand your technical breadth by working across distributed systems, cloud-native infrastructure, and optionally user-facing observability experiences.

Benefits

  • 401(k) eligibility
  • various paid time off benefits, such as PTO and parental leave
  • medical benefits
  • financial benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service