Senior Software Engineer - Databases, SRE | USA | Remote

Grafana Labs
$154,445 - $185,334Remote

About The Position

Grafana Labs is seeking a Senior Software Engineer - SRE to support its highest-value customers by enhancing the reliability of its Cloud databases (Mimir, Loki, Tempo, and Pyroscope). These databases are offered as a SaaS product across AWS, GCP, and Azure. The SRE team is integrated into the Mimir and Loki squads, focusing on delivering exceptional reliability for high-SLA customers. This role requires an engineer who can bridge customer needs, production systems, and product engineering. The company is 100% remote, with a global team of over 1,600 members across 40+ countries. They are backed by prominent investors and are committed to open source principles, collaboration, and innovation. Grafana Labs encourages applicants to apply even if they don't meet every requirement, highlighting this as a potential career-defining opportunity. This specific role is remote and targets applicants in US timezones.

Requirements

  • 6+ years engineering experience, with at least 3 years in SRE/CRE/production engineering
  • Strong preference for those with formal customer reliability engineering experience
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • Experience with Linux operating systems internals
  • Some knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high-quality PIRs
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team with a strong sense of autonomy and self-direction
  • Ability to partner deeply with product engineering teams
  • Intellectually curious, defaults to transparency, high bias towards action, and kind

Responsibilities

  • Partner closely with product engineering squads (embedded model)
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensuring customers meet SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serve as a primary escalation point and on-call for relevant incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews
  • Influence feature design to ensure production scalability and operability
  • Build automation to eliminate toil
  • Improve alert quality and reduce noisy escalations
  • Participate in on-call rotations, working with global counterparts for balanced coverage
  • Reviewing and creating SLOs, proactively investigating ways to reduce budget burn
  • Improve observability of customers within their environments
  • Design and implement solutions to ensure reliability and scalability of environments
  • Develop fault-tolerant design patterns
  • Collaborate with Engineering Leaders to define and influence product strategy, roadmaps, and technical designs
  • Participate in PR review and collaborate with other engineers on Design Docs
  • Teach others about Site Reliability Engineering and communicate best practices
  • Participate in Incident Response, including investigation, resolution, PIR, and customer communication

Benefits

  • Restricted Stock Units (RSUs)
  • 30 days of annual leave
  • 3 Grafana Shutdown Days
  • Modern AI coding assistants with company-funded usage budget
  • Access to frontier models (e.g., GPT-Codex 5/3, Claude Opus 4.6, Gemini 3 Pro)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service