The Attribution-Compression Frontier in Retrieval-Augmented Generation
| Source: arXiv AI
Tags: RAG, citation-attribution, context-compression, RECOMP, ASQA, EMNLP
Abstractive context compression in RAG achieves 0.86 citation precision against summaries it creates but only 0.12 against original source spans — a 7x attribution gap that undermines citation faithfulness in any RAG system relying on compressed context for cited answers.
Details
RAG systems that compress retrieved context before passing it to a generator face a hidden tradeoff: compression can improve efficiency but degrades the ability to verify that generated claims are grounded in source documents. This paper quantifies that tradeoff across five compression strategies (reranking, extractive selection, abstractive summarization, token pruning, and an extract-cluster-rewrite pipeline) on the ASQA and QASPER benchmarks. At a nominal 0.25 budget on ASQA, a RECOMP-style abstractive compressor achieves 0.86 precision when citations are checked against its own summaries but only 0.12 when checked against original source spans — a stark 7x gap. Extractive selection maintains more attributable grounding (0.43–0.49 precision) but at the cost of answer quality as context shrinks. A 200-question audit using TRUE T5-XXL finds emitted-grounded gaps under both fixed and recomputed source mappings. The authors note that all evaluations rely on NLI-based metrics without human calibration, so absolute numbers may shift under human review. Accepted to GroundLM 2026 Workshop at EMNLP. For enterprise RAG builders, the key takeaway is that citation precision against summaries is not a valid proxy for citation faithfulness — attribution must be evaluated against original source spans.