Head of Platform Reliability

Cielo Projects•New York, NY
•$175,000 - $230,000•Onsite

About The Position

Quant is seeking a Head of Platform Reliability to join their team. This role will be responsible for the reliability of the production environment for The Clearing House's On-Chain Money Initiative, which will allow US financial institutions to clear and settle tokenized deposits on chain. The responsibilities include managing the full production environment (OP Stack, a private Besu L1, Paladin nodes, the Overledger Gateway, the Settlement Bridge and the Audit Store), defining service levels and error-budget policy, and having the authority to stop a release when the budget is spent. The role also involves building and leading a team to run the environment on a three-shift rotation designed and participated in by the Head of Platform Reliability. The company notes that chain platforms fail differently, with state being expensive, restarts not being a strategy, and much of what appears to be infrastructure being protocol behavior.

Requirements

  • Experience leading site reliability or production engineering where failure had consequences.
  • Experience running error budgets in practice, including using one to stop a release and defending that decision.
  • Honest experience of 24x7, multi-shift coverage, understanding what causes burnout and what constitutes a livable rotation.
  • Proficiency with Kubernetes and Terraform.
  • Experience debugging an observability stack, not just configuring it.

Nice To Haves

  • Experience with blockchain nodes in production, particularly Besu or another Ethereum client.
  • Experience in banking, payments, or a regulated environment where an auditor requested proof of something.

Responsibilities

  • Own the reliability of the full production environment (OP Stack, a private Besu L1, Paladin nodes, the Overledger Gateway, the Settlement Bridge and the Audit Store).
  • Define and manage service levels and error-budget policy.
  • Exercise authority to stop a release when the error budget is spent.
  • Build and lead the team that runs the production environment.
  • Design and participate in a three-shift rotation for team coverage.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service