Senior Site Reliability Engineer

Endear,
$140,000 - $180,000Remote

About The Position

At Endear, we’re building a modern CRM for retail teams—starting with the frontline. Our software helps sales associates have more personal, effective customer conversations through AI-powered tools that drive measurable revenue. As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate. We’re hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden—not simply respond to pages. You will own the work of making Endear’s systems more observable, reliable, and scalable. You’ll investigate recurring incidents, drive root-cause fixes, improve alerting and incident playbooks, and partner with engineers on database, queue, and event-processing reliability. You will also help establish a nearshore triage layer for routine, well-understood issues, so product engineers can spend more time building product and less time on operational interruptions.

Requirements

  • Deep hands-on experience with Kubernetes, production databases, and event-driven systems.
  • Operated and improved high-volume production systems with meaningful reliability, performance, and data-scale requirements.
  • Enjoy finding root causes, fixing repeat incidents, and building tooling that makes engineers’ lives easier.
  • Experience with observability, alerting, incident response, capacity planning, and operational runbooks.
  • Can work effectively as a senior IC: owning complex technical work directly while coordinating across teams.
  • Comfortable in a lean environment where priorities move quickly and you will need to make practical trade-offs.
  • Bring experience from a scaling, mid-size company rather than only an early-stage startup or hyperscaler environment.

Nice To Haves

  • GCP or ClickHouse experience

Responsibilities

  • Build a clear view of Endear’s highest-impact reliability risks, recurring incidents, and on-call pain points.
  • Establish a prioritized reliability backlog and drive root-cause fixes for the most important issues.
  • Improve alert quality, severity definitions, escalation paths, and runbooks for common incidents.
  • Strengthen observability, queue health, database capacity planning, and operational readiness ahead of peak retail periods.
  • Create the foundation for a nearshore triage process for low-priority, repeatable issues.
  • Make on-call materially quieter and less disruptive for product engineers.
  • Build scalable systems for observability, alerting, incident response, database reliability, and queue/event-processing health.
  • Own and improve the nearshore triage relationship, playbooks, and escalation process.
  • Help establish the technical roadmap and future resourcing plan for Endear’s broader platform and reliability function.

Benefits

  • Base salary: $140,000-180,000
  • Fully remote, U.S.-based role
  • Comprehensive healthcare, including medical, dental, and vision
  • 401(k) plan
  • Monthly stipend for co-working and home-office setup
  • Flexible PTO and unlimited vacation
  • Opportunity to build Endear’s first dedicated reliability function from the ground up
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service