Stellar - Director of SRE

deCircle•New York, NY

About The Position

The Stellar Development Foundation (SDF) is seeking a Director of Site Reliability Engineering (SRE) to lead its SRE function and shape how engineering teams own, operate, and improve production services. This senior engineering leadership position reports directly to the CTO. The role involves leading a small, high-leverage SRE team while defining the broader SRE vision, operating model, and reliability culture across engineering. The focus is on creating the infrastructure, frameworks, tooling, standards, and observability practices that enable engineering teams to operate their services reliably and independently, rather than the SRE team being the operational owner of every production system. The position combines hands-on technical judgment with organizational leadership to improve reliability, infrastructure maturity, and developer productivity without introducing unnecessary process or complexity.

Requirements

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, cloud infrastructure, production operations, or closely related areas.
  • 5+ years of leadership experience, managing or formally developing SRE, infrastructure, platform, or reliability engineers.
  • Strong experience defining team charters and operating models.
  • Strong experience defining infrastructure and reliability roadmaps.
  • Strong experience defining engineering standards and practices.
  • Strong experience defining success metrics and operational maturity frameworks.
  • Deep technical judgment across distributed systems, cloud infrastructure, production operations, automation, reliability engineering, and operational risk.
  • Practical experience with AWS, GCP, or comparable cloud platforms.
  • Practical experience with Kubernetes and container orchestration.
  • Practical experience with infrastructure-as-code and declarative infrastructure.
  • Practical experience with CI/CD and deployment safety.
  • Practical experience with observability, logging, and monitoring.
  • Practical experience with SLOs and SLIs.
  • Practical experience with incident response and postmortems.
  • Practical experience with on-call systems and operational readiness.
  • Experience helping application or product engineering teams take greater ownership of production systems.
  • Understanding of how to balance developer velocity, reliability, and operational responsibility.
  • Pragmatic about tooling and comfortable deciding when to build, buy, adapt, simplify, or retire infrastructure.
  • Comfortable operating in a lean engineering organization where influence comes from technical credibility, judgment, and execution.
  • Effective communication skills with the CTO and other senior engineering leaders.

Nice To Haves

  • Leading SRE, Platform, or Infrastructure teams in lean, high-agency organizations.
  • Supporting globally distributed engineering teams and 24/7 production environments.
  • Building self-service infrastructure and paved paths.
  • Improving developer productivity through automation and toil reduction.
  • Infrastructure security, secrets management, and cloud access controls.
  • Experience in financial services or regulated environments.
  • Experience with blockchain, crypto, or Web3 infrastructure.
  • Vendor and infrastructure platform evaluation.
  • Applying AI-assisted or agentic systems to infrastructure, operations, observability, or developer workflows.

Responsibilities

  • Lead, coach, and develop a distributed SRE team, establishing its charter, priorities, operating model, and measures of success.
  • Define and roll out a Service Ownership & Maturity Framework, establishing appropriate reliability and operational standards based on service criticality.
  • Own and evolve core engineering infrastructure including cloud infrastructure, Kubernetes, CI/CD, observability, secrets management, GitHub workflows, and infrastructure-as-code.
  • Help engineering teams become stronger owners of their production services through improved dashboards, runbooks, alerting, escalation paths, and operational readiness.
  • Improve deployment automation, resilience, self-healing systems, disaster recovery, and service reliability, prioritizing based on operational risk and impact.
  • Evolve incident response, postmortems, escalation, and on-call practices across a geographically distributed engineering organization.
  • Build paved paths and self-service infrastructure to reduce engineering toil and cognitive load, enabling faster shipping without compromising reliability.
  • Collaborate with Security, Compliance, Legal, Finance, Procurement, and Corporate IT on cloud infrastructure, access management, vendors, and operational controls.
  • Pragmatically explore AI-assisted and agentic workflows to improve infrastructure operations, observability, developer productivity, and service ownership.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service