Senior Site Reliability Engineer

PlenfulSan Francisco, CA
Hybrid

About The Position

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time. You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

Requirements

  • 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
  • Operated and debugged distributed systems in production.
  • Hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
  • Defined and worked with SLOs, SLIs, and error budgets.
  • Familiarity with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
  • Can write code or scripts (Python, Bash, etc.) for automation and tooling.
  • Think in systems and reason clearly about failure modes.

Nice To Haves

  • Experience in high-growth or high-scale environments.
  • Background in regulated industries like healthcare or fintech.
  • Experience with ClickHouse or analytical systems at scale.
  • Familiarity with chaos engineering or load testing.
  • Exposure to ML infrastructure or data platforms.

Responsibilities

  • Define and implement SLIs, SLOs, and error budgets across core services.
  • Own production system health: uptime, latency, and availability targets.
  • Improve system resilience through proactive reliability work.
  • Find and mitigate single points of failure across distributed systems.
  • Take part in and improve on-call rotations and incident response.
  • Lead incident triage, mitigation, and resolution in real time.
  • Run blameless postmortems and follow through on action items.
  • Build tooling and automation to cut MTTR (Mean Time to Recovery).
  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
  • Improve signal quality to cut noise and alert fatigue.
  • Build dashboards and alerts that reflect real system health and user impact.
  • Use observability data to drive performance and reliability improvements.
  • Analyze system performance under load and find bottlenecks.
  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
  • Partner with engineering teams to improve system efficiency and scaling behavior.
  • Build automation that eliminates repetitive operational work.
  • Improve deployment safety through reliability checks and safeguards.
  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
  • Build tools for incident response, debugging, and capacity planning.
  • Partner with security and compliance to keep systems meeting operational standards.
  • Support audit readiness and reliability-related compliance requirements (Vanta).
  • Integrate monitoring and alerting into security and SIEM workflows.
  • Help mature operational practices across engineering.

Benefits

  • Full medical, dental, and vision insurance for you and participation for your family
  • 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
  • Equity — Every full-time employee shares in our success
  • Unlimited PTO — Take the time you need, when you need it
  • Daily Lunch Stipend — $100/week to cover your midday meals
  • Wellness Stipend — $100/month to support your health and well-being
  • Commuter Benefits — $100/month for SF and NYC-based employees
  • Parental Leave — Paid leave to support growing families
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service