Site Reliability Engineer, AI Observability

Appnovation Technologies•Austin, TX

About The Position

Appnovation is seeking a Site Reliability Engineer to manage a shared observability platform for LLM-based applications for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, utilizing ClickHouse for analytics, PostgreSQL for metadata, and Redis for the ingestion queue, all deployed via Argo CD. This role involves building monitoring, alerting, and service levels from scratch, as the platform currently lacks these features. Infrastructure is managed as code using Kubernetes manifests, Helm values, and Argo CD applications. The engineer will also be responsible for creating runbooks to ensure incident survivability and managing the onboarding and support process for internal teams dependent on the platform.

Requirements

  • Hands-on experience with Kubernetes on AWS (managed EKS), using Helm values and Argo CD applications.
  • Hands-on experience running ClickHouse in production, including replication, Keeper quorum, shard/replica topology, and backup/restore, preferably via an operator.
  • Experience building monitoring and alerting from scratch, including service levels reflecting user experience.
  • Experience safely upgrading self-hosted software, including schema migrations, lower environment rehearsals, and rollback plans.
  • Proficiency in PostgreSQL and Redis operations for debugging metadata-store and queue issues, including backpressure and worker drain.
  • Strong operational writing skills for runbooks, SOPs, and post-incident reviews.
  • Ability to work effectively within a client team and build trust quickly.
  • Prior experience in consulting.

Nice To Haves

  • Observability engineering, including OpenTelemetry Collector pipelines, alerting design, and Grafana dashboards.
  • OIDC or enterprise SSO integration with a corporate identity provider.
  • GitHub Actions for plan and apply pipelines with approval gates.
  • Experience running LLM observability tools such as Langfuse, LangSmith, or Arize Phoenix.
  • Experience in pharma, life sciences, or another regulated industry.
  • Prior experience and connections in the Life Sciences industry.

Responsibilities

  • Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, ingestion workers, and Kubernetes, including ingest backpressure.
  • Manage ClickHouse operations under an operator model, focusing on Keeper quorum, replication, shard/replica topology, and S3 storage tiering.
  • Rehearse, time, document, and test backup and restore procedures to ensure system recoverability.
  • Plan and rehearse safe upgrades in a lower environment, including a rollback plan for partially completed migrations.
  • Build monitoring, alerting, and service levels from scratch to proactively identify issues.
  • Maintain Kubernetes manifests, Helm values, and Argo CD applications within the delivery pipeline.
  • Write and maintain runbooks and SOPs for incident resolution by colleagues.
  • Manage the onboarding and support process for internal teams and triage incoming issues.
  • Automate recurring operational tasks.
  • Collaborate with the platform engineer on adopting and configuring platform features, managing migration and rollback.

Benefits

  • Accommodations are available upon request throughout the recruitment process.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service