Platform Reliability Technical Lead

Cielo ProjectsNew York, NY
Onsite

About The Position

When the engineer on shift cannot work out why the chain is doing what it is doing, they call you. This is the deepest technical seat on the team. Final escalation across the whole production estate, and the person who sets how the team works: runbooks that are current, alerting that fires on something real, postmortems that change something. You would work closely with the protocol engineers, because the hard problems here tend to be protocol behavior rather than infrastructure, and they get solved faster when both sides are on the same call. Very little of what you find in production will have a precedent you can look up. This role exists so that someone excellent can stay technical rather than move into management.

Requirements

  • Serious depth in reliability or production engineering.
  • You are already the person others escalate to.
  • Kubernetes, Terraform and Helm at the level where you debug them.
  • Experience with distributed systems where state matters and restarting is not a strategy.
  • Ability to document platform operations.

Nice To Haves

  • Blockchain node operations.
  • Experience reading consensus client source code to explain production behavior.

Responsibilities

  • Final escalation across the whole production estate.
  • Set how the team works: runbooks that are current, alerting that fires on something real, postmortems that change something.
  • Work closely with protocol engineers to solve hard problems related to protocol behavior.
  • Debug production issues that may not have a precedent.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service