CloudZero is growing rapidly, with an expanding customer base and increasingly complex data challenges. The platform is scaling to match this growth. A key initiative is the implementation of real-time ingestion on Kafka, involving multiple engineering teams and presenting significant operational demands. Currently, no single entity owns the end-to-end reliability of this critical path. This role will initially focus on owning this reliability. As a Senior Site Reliability Engineer, you will act as a force multiplier for the engineering organization, taking ownership of the reliability, performance, and observability of systems essential to all teams. You will empower teams to deliver features that enable customers to understand and optimize their cloud spending. This position involves substantial infrastructure work at a significant scale, moving beyond simple ticket resolution or console operations. CloudZero processes billions of events daily across AWS, Azure, and GCP. Customers depend on accurate, real-time cost data for critical business decisions, making system stability paramount. The platform is built on a unique serverless architecture without EC2 instances or containers, requiring infrastructure that scales gracefully, fails predictably, and self-heals automatically. The focus is on engineering solutions for reliability rather than reactive firefighting, as there are no Kubernetes clusters or broker fleets to tune. On-call responsibilities are manageable: Nimbus handles a weekly rotation for shared infrastructure, while feature teams are responsible for their own services. This role is ideal for individuals who excel at solving complex operational problems, have a strong commitment to reliability and performance, and desire to see their work have a direct and measurable customer impact.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed