Lead Site Reliability Engineer

Stuut•San Francisco, CA

2d•$200,000 - $275,000

About The Position

Stuut is transforming accounts receivable for B2B companies, making collections smarter and faster. Our platform is gaining traction with finance teams across various sectors. We are seeking a Lead Site Reliability Engineer to drive the strategy, architecture, and execution of reliability, scalability, and operational excellence across our platform. This role involves building and scaling systems to ensure Stuut remains highly available, performant, and resilient as the company grows. The Lead SRE will define SLOs and reliability standards, harden infrastructure, improve observability, and guide teams through incident response and postmortems. This is a hands-on technical leadership role for an engineer skilled in designing reliable distributed systems, influencing engineering practices, and leading high-impact reliability initiatives.

Requirements

7+ years of experience in site reliability engineering, infrastructure engineering, or backend software engineering.
Designed and operated highly available, production-grade systems supporting rapid product iteration.
Fluent in Python and/or TypeScript, and comfortable building automation and tooling to support reliability goals.
Deep experience with AWS, Kubernetes (EKS), Docker, and cloud-native architectures.
Implemented and evolved observability stacks (metrics, logs, traces) and know how to create high-signal alerting.
Understand how to design, measure, and enforce SLOs, SLIs, and error budgets.
Supported systems built with modern stacks such as FastAPI, Vue.js, PostgreSQL (RDS), and event-driven architectures.
Improved reliability and operational maturity in environments using CI/CD pipelines, infrastructure as code, and modern deployment workflows.
Can balance reliability, velocity, and cost — making pragmatic tradeoffs that serve customers and the business.
Enjoy collaborating across Product, Backend, Frontend, and Infrastructure teams to improve system health.
Thrive in a role that blends deep technical execution, system design, and leadership influence in a fast-moving environment.

Responsibilities

Set the Reliability Strategy: define the long-term vision for site reliability, including SLOs/SLIs, error budgets, availability targets, and operational standards.
Build & Scale Reliable Infrastructure: architect and maintain resilient, scalable cloud infrastructure across AWS and Kubernetes, ensuring systems are secure, fault-tolerant, and cost-effective.
Own Observability & Monitoring: design and evolve monitoring, alerting, and logging systems that provide clear, actionable signals across services and environments.
Lead Incident Response & Postmortems: own incident management practices, lead major incident response, and drive blameless postmortems that result in meaningful system improvements.
Improve System Resilience: identify reliability risks and lead efforts around redundancy, failover, capacity planning, and graceful degradation.
Optimize CI/CD & Deployment Reliability: partner with engineering teams to ensure deployments are safe, observable, and reversible; improve rollout strategies and reduce operational risk.
Partner with Product & Engineering Teams: collaborate early in the development lifecycle to influence system design, scalability, and reliability tradeoffs.
Reduce Toil & Improve Developer Experience: automate operational tasks, improve runbooks, and build tooling that reduces manual work and accelerates safe execution.
Drive Root Cause Resolution: guide teams through deep debugging of reliability issues, ensuring fixes address underlying causes rather than symptoms.
Influence Reliability Culture: promote reliability-first thinking, strong operational hygiene, and shared ownership of production systems across engineering.
Mentor & Level Up the Team: coach engineers on reliability principles, incident handling, infrastructure design, and operational best practices.