Senior Software Engineer, Observability Delivery

GitHub, Inc.•UNAVAILABLE, UNAVAILABLE
•Remote

About The Position

GitHub’s Observability Delivery team builds and operates the company-wide pipelines for logs, metrics, traces, and exceptions that engineers rely on to monitor and diagnose their services. These systems handle telemetry from across GitHub at high scale. Keeping them reliable and efficient as demand grows is a constant engineering challenge. We’re looking for a Senior Software Engineer who enjoys solving challenges across software and infrastructure. You’ll write and operate services, deploy and configure observability tools, and make architectural decisions about how telemetry moves through our systems. You’ll lead work to improve pipeline capacity, reliability, and efficiency while partnering with service teams throughout GitHub. If you enjoy making critical, large-scale infrastructure dependable and easier to operate, we’d like to hear from you.

Requirements

  • 6+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Associate's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Bachelor's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 4+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 2+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field OR equivalent experience.
  • 3+ years building and operating production services or infrastructure in a cloud or large-scale distributed environment.
  • 2+ years deploying, configuring, or troubleshooting production workloads with Kubernetes or comparable container orchestration.
  • 2+ years developing production software in Go, Ruby, or a comparable general-purpose language.

Nice To Haves

  • Experience building or operating shared platforms used by other engineering teams.
  • Experience with logging, metrics, or distributed tracing systems, including OpenTelemetry concepts or tooling.
  • Experience operating telemetry pipelines or platforms such as Datadog or comparable tools.
  • Experience with Azure or another major cloud platform, alongside hybrid or self-managed infrastructure.
  • Familiarity with networking, service connectivity, and distributed-systems operations.
  • Experience leading ambiguous technical work across teams and mentoring engineers through design and code review.

Responsibilities

  • Design, build, and operate high-scale pipelines for logs, metrics, traces, and exceptions, balancing capacity, reliability, performance, and cost.
  • Lead cross-team work to identify and remove scaling bottlenecks, improve how telemetry is collected and processed, and safely evolve critical production infrastructure.
  • Write and maintain production software, primarily in Go and Ruby, while configuring and integrating open-source and commercial observability tools.
  • Work across cloud infrastructure, Kubernetes, virtual machines, networking, and service connectivity to make distributed systems dependable and operable.
  • Partner with service teams to understand their observability needs and improve the tools they use to diagnose and maintain their services.
  • Own system health through monitoring, incident response, on-call participation, and improvements informed by operational experience.
  • Provide technical leadership through design proposals, reviews, mentoring, and collaboration across teams.

Benefits

  • competitive pay
  • generous learning and growth opportunities
  • excellent benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service