Staff Site Reliability Engineer

FinalsiteGlastonbury, CT
Remote

About The Position

As a Staff Site Reliability Engineer, you'll focus on the platform behind Finalsite's Composer CMS, while also working cross-team to shape platform-wide standards and integrations. You'll help set the technical direction for how we build, scale, and operate our platform in a GCP-primary, multi-cloud environment. You'll be the person other engineers come to when a system needs to be rethought, not just repaired, and you'll play a key role in growing the senior engineers around you. This is a role for someone who wants to know exactly how the systems work, how they will fail, and how to build the guardrails that keep everyone else from finding out the hard way.

Requirements

  • Expert-level knowledge of GCP and infrastructure broadly, comfortable operating across the full stack rather than one layer of it.
  • Expertise in IaC, including Terraform and Terragrunt.
  • Experience with GitLab (or similar) and CI/CD pipeline design and operation.
  • Expertise in Kubernetes and containerization, including Helm.
  • Strong experience with network architecture, including routing and connectivity within GCP and across cloud providers.
  • Strong experience with edge security and traffic management, using solutions like Cloudflare.
  • Deep familiarity with observability tooling and monitoring strategy.
  • Real experience leading incident management and response, not just participating in it.
  • A track record of operating at a staff level, including mentoring senior engineers.
  • Experience architecting highly available, fault-tolerant systems.
  • Can read, understand, and write basic code when needed, across languages and runtimes such as Ruby/Rails, Java, Python, for example.
  • Comfortable in application code to spot reliability, performance, or architecture issues.

Nice To Haves

  • Experience with cloud cost allocation and control.
  • Experience with capacity planning and forecasting for infrastructure at scale.
  • Familiarity with AWS and/or Azure ( GCP-primary ).

Responsibilities

  • Guide our cloud architecture strategy with a focus on scalability, maintainability, and cost efficiency.
  • Lead capacity planning so we scale ahead of demand instead of reacting to it.
  • Own the health of our Kubernetes (GKE) platform, from cluster architecture to workload reliability at scale.
  • Design and evolve our network architecture, including global routing, connectivity, and edge strategy with Cloudflare.
  • Drive IaC standards across the team, building reusable Terraform/Terragrunt modules that other teams can adopt without reinventing them.
  • Build and mature our observability practice, defining what we monitor, how we alert, and how we define and hold ourselves to SLOs.
  • Lead disaster recovery planning for the platforms you own, including backup design, failover planning, and clear recovery objectives (RTO/RPO).
  • Architect highly available, fault-tolerant systems.
  • Bring a security-first mindset to everything you build, treating it as a design input from day one, not a review gate at the end.
  • Keep a close eye on cloud cost and help teams make smart tradeoffs between performance, resilience, and spend.
  • Read, understand, and write basic code when needed, across languages and runtimes such as Ruby/Rails, Java, Python, for example.
  • Mentor senior engineers, helping them grow their technical judgment and take on bigger calls of their own.
  • Build tools and patterns that let other teams move independently, turning one-off solutions into lasting, self-serve practice, so SRE isn't the bottleneck standing between people and the answer.
  • Represent infrastructure and reliability concerns in planning conversations, weighing in early enough to shape decisions, not just implement them.
  • Champion strong change management practices, including peer review, staged rollouts, and go/no-go gates for high-risk changes.
  • Lead response for high-severity incidents, participate in our on-call rotation, and turn every incident into a lasting improvement.

Benefits

  • Competitive benefits
  • Professional development opportunities
  • Collaborative culture built on partnership and purpose
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service