The Senior Site Reliability Engineer is accountable for how Stratus behaves in production. Reporting to the Senior Director, Platform Engineering, this role brings genuine SRE discipline to a platform that MEP contractors run their fabrication shops on — where downtime does not mean a degraded experience, it means work stops on a job site. Stratus is a ~50-person, primarily remote, Series B software company growing quickly. Unrelenting Reliability is one of our company values, and this is the role that operationalizes it. You will define what reliable means in numbers, instrument the system so we can see it, and close the loop from production signal back into engineering priority. This is an engineering role with a production mandate: you will write code, tune queries, build alerting, run load tests, and lead incidents — and you will be measured on customer-visible availability and latency, not on tickets closed. The load-bearing problems on your plate are: (1) establishing service level objectives and an error budget the whole engineering organization operates against, with the measurement infrastructure to back them; (2) building the detection and response capability — high-signal alerting, clear on-call and escalation paths, per-system incident ownership, and well-exercised recovery paths — so that production problems are caught and resolved fast; and (3) building the performance and capacity engineering practice for our data and messaging layers, with headroom measured rather than assumed. You will work across every engineering pod, with our platform and security functions, and with the customer-facing teams who see reliability problems first. The right candidate is comfortable being the person who says a number out loud and defends it, and is drawn to a place where the reliability practice is being built rather than maintained.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed