Appnovation is seeking a Site Reliability Engineer to manage a shared observability platform for LLM-based applications for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, utilizing ClickHouse for analytics, PostgreSQL for metadata, and Redis for the ingestion queue, all deployed via Argo CD. This role involves building monitoring, alerting, and service levels from scratch, as the platform currently lacks these features. Infrastructure is managed as code using Kubernetes manifests, Helm values, and Argo CD applications. The engineer will also be responsible for creating runbooks to ensure incident survivability and managing the onboarding and support process for internal teams dependent on the platform.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed