Senior Platform Reliability Engineer

Grow TherapySan Francisco, CA
Hybrid

About The Position

Grow Therapy is seeking a Senior Platform Reliability Engineer to establish and scale reliability as a core capability within the organization. This role will operate across departments, influencing how reliability is understood, measured, and integrated into the developer experience. The engineer will collaborate with platform and product engineering teams to set standards for observability, SLOs/SLAs, and incident response, and translate these into self-service tools and best practices. This is a high-impact, autonomous position focused on driving both cultural and technical change to enable teams to build and operate reliable systems at scale.

Requirements

  • 6+ years of experience operating and improving the reliability of production systems at scale.
  • Hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform.
  • Deep understanding of reliability principles, including SLOs/SLAs, error budgets, and improving reliability through measurement and iteration.
  • Experience with modern observability tooling (e.g., DataDog) and building actionable monitoring systems across metrics, logs, and traces.
  • Systems thinking ability to identify patterns and design scalable solutions.
  • Impact-oriented focus on outcomes and improving real reliability.
  • Strong communication and influencing skills to drive change across teams without direct authority.
  • Self-directed and comfortable defining problems, proposing solutions, and executing independently in ambiguous environments.
  • Team player with strong collaboration, empathy, and mentoring skills.

Nice To Haves

  • Experience introducing or scaling reliability practices in a growing organization.
  • Experience building internal tooling or platforms used by multiple teams.
  • Experience designing service-level scorecards or compliance/reporting systems.
  • Experience with both SaaS (e.g., DataDog) and self-managed observability stacks.
  • Previous experience as a product engineer, bringing empathy for developer experience.
  • Experience with database reliability and performance (e.g., PostgreSQL).

Responsibilities

  • Define and scale reliability as a discipline by establishing frameworks for SLOs/SLAs, error budgets, and operational readiness.
  • Improve observability and measurement by identifying gaps in metrics, logging, and tracing, ensuring services are measurable and debuggable.
  • Evolve incident response practices, from detection to post-incident learning, and help teams build sustainable on-call and escalation patterns.
  • Enable self-service reliability by partnering with the platform team to build tooling and abstractions that facilitate adoption and compliance with reliability standards.
  • Drive adoption across teams by educating, influencing, and guiding engineering teams through clear standards, strong communication, and developer-friendly systems.

Benefits

  • Comprehensive medical, dental, vision, life, and disability coverage.
  • No cost access to therapy through the Grow platform (available to US employees).
  • Retirement savings programs and equity opportunities.
  • Flexible time off, company paid holidays, and a full company-wide Winter Break.
  • Up to 18 weeks of paid parental leave and a new child stipend.
  • Dedicated weekly flexible time for self-care (Mental Health Mornings/Afternoons).
  • Annual stipend for personal wellbeing and professional growth.
  • Pre-tax commuter benefits.
  • Support for home workspace and meal benefits.
  • Variety of wellbeing benefits, including wellness memberships, virtual care, pet insurance discounts, and global travel assistance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service