About The Position

Wave is seeking an Observability Engineer to join the newly formed Performance & Observability team. This role is crucial in helping engineers understand, operate, debug, and improve the systems that power payments for millions of users across Africa. The Observability Engineer will own and evolve Wave’s observability platform across various technologies including Python backend, GraphQL API, Postgres and CockroachDB databases, Kubernetes workloads, cloud infrastructure, and on-premises environments. The work will directly contribute to making it easier for product, database, infrastructure, and security teams to detect problems early, understand system behavior, resolve incidents quickly, and make informed reliability and performance decisions. The role reports to the Director of Platform and collaborates closely with infrastructure, database, and product engineering teams.

Requirements

  • 5+ years of experience in observability, SRE, platform engineering, infrastructure engineering, backend engineering, or production systems engineering.
  • Deep understanding of metrics, logging, tracing, profiling, alerting, dashboards, service-level indicators, and incident response workflows at scale.
  • Experience building internal tools, libraries, automation, or platforms used by other engineers.
  • Excellent communication and collaboration skills. This role succeeds by helping other engineers build, operate, and debug better systems.
  • Pragmatic judgment about when to improve tooling, when to simplify, and when to avoid unnecessary complexity.
  • Experience learning, analyzing, and working on other people’s code.
  • Proficiency in at least one backend language, preferably Python.
  • Experience with observability tools such as Prometheus, Grafana, Datadog, OpenTelemetry, Jaeger, Tempo, Loki, Honeycomb, Sentry, or similar systems.
  • Experience with at least some our stack, Postgres, CockroachDB, Redis, GraphQL, or Kubernetes.
  • Experience with OpenTelemetry instrumentation and collector configuration at scale.

Nice To Haves

  • Care a lot about working on software whose mission you can believe in.
  • Have a bias for action. You see a problem, you fix a problem. You get buy-in for your solutions and keep work moving.
  • Approach your work with a growth mindset and use your skills and experience to mentor less-experienced engineers.
  • Are excited to build world-class infrastructure that powers economic opportunity for an entire continent.

Responsibilities

  • Improve understanding of production behavior across application code, GraphQL APIs, databases, caches, Kubernetes workloads, cloud infrastructure, async jobs, and on-premises environments.
  • Partner with product and platform teams to define meaningful service-level indicators, reduce alert fatigue, and ensure alerts are actionable, reliable, and tied to user impact.
  • Build internal tooling and self-service workflows that help engineers instrument services, investigate incidents, analyze performance, and understand dependencies.
  • Help teams identify reliability, latency, capacity, and cost issues before they become user-facing incidents.
  • Establish and implement observability standards, documentation, and training materials that scale across Wave’s engineering organization.
  • Operate and improve our observability platform (e.g., Datadog, Honeycomb, Prometheus, Grafana, OpenTelemetry) while controlling costs as data volume grows.
  • Ownership, management and evolution of observability tools: Datadog, Honeycomb, Sentry, and Pyroscope.
  • Design of SLOs and support product teams in implementing them.
  • Support all teams in improving their alert quality.
  • Ownership of observability modules and code, ensuring consistent naming conventions across all our systems.
  • Early detection of regressions in our code (performance, reliability, or other regressions impacting users).
  • CPU profiling in production.
  • SLO support for product teams.
  • Scaling our Observability platform to support 4x our current volume while keeping costs in check.

Benefits

  • We foster autonomy for our employees. You'll own your projects at every stage, from understanding the problem to monitoring your solution in production.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service