Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering
| Source: arXiv AI
Tags: knowledge graphs, multi-hop QA, reinforcement learning, Qwen3, biomedical NLP
A training pipeline combining context-augmented KG supervision, an adaptive repair stage that achieves 100% one-hop accuracy, and RL initialization consistently improves multi-hop question answering with Qwen3-14B across disease-specific knowledge graphs.
Details
Multi-hop question answering over knowledge graphs requires models to chain across multiple fact triples. Training LLMs on isolated head-relation-tail triples often fails because the model lacks surrounding context. This paper proposes a three-stage framework evaluated on medical KGs for gastroparesis and diabetes. First, context-augmented training: for each target KG triple, supporting triples from the same source text chunk are attached to form a context graph (CG). Two variants are trained: KGModel (isolated triples only) and CGModel (triple + context). The CG-grounded approach consistently improves multi-hop performance. Second, an adaptive repair pipeline: LLM-judged, history-aware identification of persistent one-hop failures, followed by targeted fine-tuning on repair examples and removal of noisy triples. This stage brings one-hop accuracy to 100% on the cleaned retained validation set. Third, RL fine-tuning initialized from the repaired SFT checkpoint using lower-hop QA items, then evaluated on harder 3-hop, 4-hop, and 5-hop tasks. RL from repaired checkpoints yields larger and more stable gains than RL from non-repaired ones.