TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
| Source: arXiv AI
Tags: RAG, AIOps, on-premise LLM, root cause analysis, vLLM, Mistral, Qwen
TriCalRAG shows RAG over labeled incident history improves on-premise LLM root cause analysis F1 by 0.10–0.27 and eliminates the near-degenerate behavior where zero-shot models predict 'anomaly' on 100% of incidents; batching scales throughput 41x and 4-bit quantization cuts latency 20% with no accuracy loss on a single RTX PRO 6000.
Details
Cloud-hosted LLMs for AIOps introduce data privacy risk, latency, and per-query cost that scale poorly with production log volumes. TriCalRAG benchmarks open-weight LLMs served locally with vLLM against a classical LSTM-based log anomaly detector (DeepLog) across four real-world log datasets: BGL, HDFS, Thunderbird, and OpenStack.\n\nTwo models (Qwen2.5-14B and Mistral-Small) were evaluated under zero-shot, few-shot, and RAG prompting with bootstrap 95% confidence intervals across three random seeds. The most important finding is calibration: zero-shot prompting drives both models toward near-degenerate behavior (predicting 'anomaly' on up to 100% of incidents on some datasets), while RAG keeps predicted-positive rates close to true class balance.\n\nMistral-Small achieves higher macro-averaged F1 (0.644 vs. 0.560 for Qwen2.5-14B) but fails calibration in more configurations (7 vs. 5 of 12) — meaning better accuracy does not guarantee predictable behavior across prompting conditions. An ablation shows batching scales throughput 41x on a single GPU and 4-bit quantization reduces latency 20% with no measurable accuracy loss.\n\nFor enterprise IT and DevOps teams evaluating on-premise AI for incident response, this benchmark provides directly applicable guidance on model selection, prompting strategy, and hardware configuration without requiring cloud API access.