About The Position

We are looking for a full-time, permanent Senior Platform & Reliability Engineer to start as soon as possible. We live remote-first, but you have the freedom to choose whether you want to work hybrid or completely on-site due to your proximity to one of our locations (Berlin, Cologne, Hamburg, Munich). As a Senior Platform & Reliability Engineer at Contabo, you take architectural ownership of our shared infrastructure services – the foundation that multiple development teams build on: API gateway and ingress (Kong, Nginx Ingress), persistent storage (Ceph, Longhorn), secrets and identity infrastructure (Vault, Keycloak), edge security and WAF (Cloudflare), and our observability stack. You'll be taking over grown, partially under-documented systems – and that's exactly where the appeal of this role lies: you work your way deep into these systems, identify and remediate known weak points, evaluate aging components with solution-agnostic build-vs-buy reasoning, and decide which legacy pieces get fixed, replaced, or retired. One of your central mandates: you design and establish an on-call process, including runbooks for platform and infrastructure incidents – where no formal process exists today. In parallel, you mature our observability practices, rolling out distributed tracing, SLOs/SLIs, and meaningful dashboards across the stack, building on our existing tooling with Prometheus, Grafana, Alloy, and OpenTelemetry. In system design reviews, you bring your strong grounding in the fundamentals – load balancing, caching, sharding, replication, consistency trade-offs – and apply them to new and existing services alike. You won't be managing a team, but you will be the technical authority multiple teams rely on: you advise across teams on shared infrastructure services, document your knowledge consistently, and actively distribute it – so that critical know-how never again depends on a single person. Success in this role means: known risks are resolved, a solid on-call process is up and running, and single points of failure are measurably reduced platform-wide. The position is remote (Germany), with hybrid or on-site work optional; occasional travel for datacenter visits and team offsites is part of the role.

Requirements

  • 7+ years of experience in platform, infrastructure, or SRE roles, ideally with end-to-end responsibility for a private cloud or IaaS platform.
  • Hands-on production experience with distributed storage systems (Ceph) and Kubernetes persistent storage (Longhorn or comparable).
  • Experience operating API gateways and ingress (Kong, Nginx Ingress, or comparable), including debugging cross-cutting concerns like CORS and rate limiting.
  • Experience with CDN/edge security, DDoS mitigation, and WAF configuration (e.g. Cloudflare), as well as firewall-rule design.
  • A strong grounding in system design fundamentals (load balancing, caching, sharding/replication, consistency models, message queues) – and the judgment to apply them to real-world trade-offs.
  • Professional fluency in English (working language).

Nice To Haves

  • Experience designing or maturing observability (tracing, metrics, SLOs/SLIs) with tools such as Prometheus, Grafana, and OpenTelemetry.
  • Experience with secrets/identity infrastructure (Vault, Keycloak), messaging systems (NATS), and building on-call processes and incident runbooks from scratch.
  • Composure working with grown, incompletely documented systems – paired with the right mix of pragmatism and perfectionism.
  • Genuine enjoyment of acting as the technical go-to person across teams, actively sharing and documenting your knowledge.
  • German language skills.
  • Certifications (CKA/CKS, Ceph training).
  • Experience with virtualization platforms (Proxmox, OpenStack).

Responsibilities

  • Take architectural ownership of shared infrastructure services including API gateway and ingress, persistent storage, secrets and identity infrastructure, edge security and WAF, and observability stack.
  • Identify and remediate weak points in existing systems.
  • Evaluate aging components and decide on their future (fix, replace, or retire).
  • Design and establish an on-call process with runbooks for platform and infrastructure incidents.
  • Mature observability practices by rolling out distributed tracing, SLOs/SLIs, and meaningful dashboards.
  • Apply system design fundamentals (load balancing, caching, sharding, replication, consistency trade-offs) to new and existing services.
  • Advise across teams on shared infrastructure services.
  • Document and distribute knowledge consistently.
  • Resolve known risks, ensure a solid on-call process is operational, and reduce single points of failure.

Benefits

  • Workation across the EU and in our summer office in Mallorca
  • Access to EGYM Wellpass and thousands of fitness and wellness facilities
  • Attractive discounts on many products and services through our corporate benefits program
  • 30 days of vacation plus additional days off on Christmas Eve and New Year’s Eve
  • An extra day off to get involved in social activities through our Volunteer Day
  • Individual professional and personal development opportunities
  • Modern, conveniently located offices across our European locations
  • Work in an international environment shaped by diversity
  • Company events and team activities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service