Site Reliability Engineer 3

Granicus India
Remote

About The Position

Granicus is seeking a Site Reliability Engineer 3 (SRE) with strong AIOps, automation, and AI proficiency to modernize reliability engineering through observability, intelligent incident response, and responsible AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build scalable, resilient platforms supporting traditional, cloud-native, and AI/ML-powered workloads. The role will also help operationalize AI-enabled SRE practices such as alert intelligence, assisted root-cause analysis, runbook automation, telemetry summarization, and governed self-healing workflows with appropriate human approval and audit controls. You will be expected to implement practical AIOps capabilities across observability, incident response, automation, and operational knowledge workflows, turning AI/ML insights into production-ready reliability improvements.

Requirements

  • 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments.
  • Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP).
  • Expert in ELK/OpenSearch, including: Log ingestion (Logstash / Beats), Elasticsearch index design, scaling, and tuning, Advanced Kibana querying and debugging, Dashboards, alerts, and observability patterns for production systems
  • Hands-on experience in logs, metrics, and tracing.
  • Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata.
  • Solid understanding of incident management, RCA, SLOs, and operational best practices.
  • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting.
  • Hands-on implementation experience with AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation.
  • Ability to implement AIOps integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.
  • Experience measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, faster MTTD/MTTR, RCA quality, automation adoption, repeat usage, and business impact linkage.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar.

Nice To Haves

  • Preferred Certifications : AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.

Responsibilities

  • End-to-end reliability for production systems: on-call, incident response, postmortems, SLO/SLI ownership
  • Building and maintaining observability pipelines (metrics, logs, traces) with AI-assisted anomaly detection layered on top and implementing AIOps pipelines for event ingestion, enrichment, deduplication, correlation, and noise reduction
  • Designing automation that reduces toil — with a clear bias toward AI-augmented runbooks over static scripts and implementing AIOps-driven remediation workflows with approvals, rollback logic, and audit trails
  • Driving alert-noise reduction using correlation/ML techniques, not just threshold tuning
  • Partnering with engineering teams to embed reliability and AI-assisted diagnostics into the SDLC
  • Demonstrated production use of LLMs/AI agents for SRE workflows — e.g., automated log triage, RCA drafting, runbook generation, or incident summarization — not just "I used Copilot to write YAML"
  • Experience building or integrating AIOps capabilities: anomaly detection, alert correlation/clustering, predictive capacity signals including hands-on implementation using observability platforms, ML-based signal processing, incident enrichment, and ChatOps/ticketing integrations
  • Working knowledge of prompt engineering for operational use cases (structured outputs, tool-use/function calling, retrieval-augmented context from runbooks/CMDB)
  • Comfort evaluating AI output critically — can articulate where an LLM's suggested fix or RCA was wrong and why, not just accept it
  • Familiarity with agentic frameworks or MCP-style tool integration (connecting LLMs to ticketing, observability, or ChatOps tools) is a strong plus
  • Understanding of the risk surface of AI-in-production: hallucination in RCA, over-automation risk, human-in-the-loop design for high-blast-radius actions
  • Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration.
  • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups.
  • Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to accelerate diagnosis while validating AI recommendations before action.
  • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting.
  • Design and maintain observability across logs, metrics, and traces, ensuring actionable monitoring and high signal-to-noise alerting.
  • Build and enhance workflows for alerting, anomaly detection, and incident enrichment to reduce noise and improve accuracy.
  • Implement AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment.
  • Develop automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback plans, and auditability.
  • Implement AIOps remediation patterns that connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps.
  • Drive improvements in system reliability, performance, scalability, and resilience through engineering-led initiatives.
  • Partner with engineering teams to improve deployment safety, operational readiness, and production stability.
  • Maintain high-quality runbooks, documentation, and knowledge bases to improve on-call effectiveness and knowledge sharing.
  • Support capacity planning, performance tuning, and SLO-based reliability practices.
  • Apply security, access control, and operational guardrails across systems and automation.
  • Own AIOps implementation from use-case definition through production rollout, including telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.

Benefits

  • Employee Resource Groups to encourage diverse voices
  • Coffee with Mark sessions – Our employees get to interact with our CEO on very important and sometimes difficult issues ranging from mental health to work-life balance and current affairs.
  • Microsoft Teams communities focused on wellness, art, furbabies, family, parenting, and more.
  • We bring in special guests from time to time to discuss issues that impact our employee population
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service