Senior CloudOps Engineer

CloudZeroBoston, MA
$130,000 - $190,000

About The Position

CloudZero is growing rapidly, with an expanding customer base and increasingly complex data challenges. The platform is scaling to match this growth. A key initiative is the implementation of real-time ingestion on Kafka, involving multiple engineering teams and presenting significant operational demands. Currently, no single entity owns the end-to-end reliability of this critical path. This role will initially focus on owning this reliability. As a Senior Site Reliability Engineer, you will act as a force multiplier for the engineering organization, taking ownership of the reliability, performance, and observability of systems essential to all teams. You will empower teams to deliver features that enable customers to understand and optimize their cloud spending. This position involves substantial infrastructure work at a significant scale, moving beyond simple ticket resolution or console operations. CloudZero processes billions of events daily across AWS, Azure, and GCP. Customers depend on accurate, real-time cost data for critical business decisions, making system stability paramount. The platform is built on a unique serverless architecture without EC2 instances or containers, requiring infrastructure that scales gracefully, fails predictably, and self-heals automatically. The focus is on engineering solutions for reliability rather than reactive firefighting, as there are no Kubernetes clusters or broker fleets to tune. On-call responsibilities are manageable: Nimbus handles a weekly rotation for shared infrastructure, while feature teams are responsible for their own services. This role is ideal for individuals who excel at solving complex operational problems, have a strong commitment to reliability and performance, and desire to see their work have a direct and measurable customer impact.

Requirements

  • Strong production Python skills as the primary language, with experience in owning, testing, and maintaining code at scale.
  • Experience defining an SLO, including deliberate decisions about what not to alert on.
  • Experience operating asynchronous, event-driven systems and understanding concepts like back-pressure, consumer lag, replay, poison messages, and partial failure (Kafka, Kinesis, SQS, Pulsar, or Step Functions).
  • Demonstrated ability to drive a reliability or platform change through a team that did not report to you and had not requested it.
  • Minimum of 5 years of experience building and operating distributed systems in AWS, with a focus on owning reliability outcomes rather than just tasks.
  • Practical experience with Infrastructure as Code using CloudFormation and SAM, or equivalent depth in Terraform or Pulumi.
  • Hands-on experience instrumenting systems in monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk.
  • Proven ability to debug production issues under pressure.
  • Interest in frontier AI models like Claude, Codex, or Gemini.
  • Preference for thoughtful, reliable system design over reactive "hero" efforts.
  • Strong documentation habits to ensure long-term team clarity and system stability.
  • Ability to clearly explain complex technical issues to non-technical stakeholders.
  • Comfort with managing multiple areas simultaneously, diving deep into critical issues, and then standardizing solutions before moving on.

Nice To Haves

  • Chaos engineering or load testing as a practice you built.
  • Internal developer portal experience (e.g., Cortex, Backstage).
  • Experience with test automation or ephemeral test environments.
  • Experience with GitHub Actions at scale.
  • Experience with LLM-backed tooling used daily by engineers.

Responsibilities

  • Own the reliability practice for CloudZero's real-time ingestion path, including SLOs that span team boundaries, failure modes in system seams, and architectural improvements based on learnings.
  • Approve shared critical paths before deployment and initiate pause discussions when error budgets are depleted.
  • Instrument systems to ensure rapid failure detection and data-driven debugging.
  • Build observability into all systems to proactively identify and address issues before customers are affected.
  • Develop reliability tooling such as load generators, fault-injection harnesses, SLO instrumentation libraries, and deployment safety checks.
  • Write production Python for shared libraries, internal services, automation, and agents, establishing coding standards for others.
  • Design and maintain CloudFormation and SAM modules for provisioning reliable and cost-efficient cloud resources.
  • Manage infrastructure end-to-end without manual console interaction.
  • Automate deployments, scaling, backups, and limit changes, replacing repetitive human tasks with systems.
  • Implement intelligent automation, focusing on solving real problems rather than automating for its own sake.
  • Evaluate existing autonomous agents in production, enhancing effective ones and retiring ineffective ones.
  • Enhance system legibility for AI tooling and human understanding by incorporating service and ownership metadata into the developer portal.
  • Collaborate with Product Engineering teams to design resilient services, review architectures for operational complexity, and build deployment pipelines for safe and rapid feature delivery.
  • Integrate SLOs and instrumentation into shared templates to promote best practices.
  • Drive adoption of reliability practices among 40+ engineers through clear communication and demonstrated value.
  • Optimize for cost and performance, serving as an example of efficient cloud usage.

Benefits

  • Collaborative, fast-moving environment
  • Work with cutting-edge technology
  • Drive meaningful outcomes
  • Opportunity to grow with a fast-scaling company
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service