Head of Platform Reliability

Cielo Projects•New York, NY
•Onsite

About The Position

Quant is seeking a Head of Platform Reliability to join their team. Quant is a leading provider of programmable money infrastructure, enabling banks to issue, move, and settle tokenized deposits around the clock. Their technology is deployed in regulated environments with central and commercial banks globally, including work on the digital pound and digital euro. The Clearing House has selected Quant to power its On-Chain Money Initiative, a new interoperable payments network for financial institutions. Quant provides the network's interoperability, orchestration, and transaction-management layer. This role will own the reliability of the full production environment for the On-Chain Money Initiative, including the OP Stack, a private Besu L1, Paladin nodes, the Overledger Gateway, the Settlement Bridge, and the Audit Store. The position involves defining service levels, error-budget policy, and having the authority to stop releases when the budget is spent. The role also includes building and leading a team to run the production environment on a three-shift rotation designed and participated in by the Head of Platform Reliability. The role requires an understanding of how chain platforms fail differently, where state is expensive, restarts are not a strategy, and protocol behavior is key.

Requirements

  • Experience leading site reliability or production engineering where failure had consequences.
  • Experience running error budgets in practice, including using one to stop a release and defending that decision.
  • Honest experience of 24x7, multi-shift coverage.
  • Understanding of what causes burnout and what does not, and the difference between a rotation on paper and one people can live with.
  • Proficiency in Kubernetes.
  • Proficiency in Terraform.
  • Experience debugging an observability stack, not just configuring it.

Nice To Haves

  • Experience with blockchain nodes in production, particularly Besu or another Ethereum client.
  • Experience in banking, payments, or a regulated environment where an auditor required proof of something.

Responsibilities

  • Own the reliability of the full production environment (OP Stack, a private Besu L1, Paladin nodes, the Overledger Gateway, the Settlement Bridge and the Audit Store).
  • Define and manage service levels and error-budget policy.
  • Exercise authority to stop a release when the error budget is spent.
  • Build and lead the team responsible for running the production environment.
  • Design and participate in a three-shift rotation for 24x7 coverage.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service