GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning
| Source: arXiv AI
Tags: RAG, multi-agent systems, graph neural networks, reinforcement learning, multimodal AI, long-horizon reasoning
GraMRAG proposes a multi-agent RAG framework that models reasoning state as a dynamic directed acyclic graph, using topology-aware policy optimization to identify critical reasoning paths—achieving state-of-the-art on complex multimodal long-horizon reasoning benchmarks.
Details
Multi-agent RAG systems for complex reasoning tasks suffer from two documented problems: inadequate retrieval depth and state blindness—they fail to track what has been explored or concluded. GraMRAG addresses both by formalizing agent reasoning as a dynamic directed acyclic graph (DAG) where nodes represent actions and edges model action-observation dependencies.\n\nThree components work together: a vision-text bridged reasoning paradigm that combines multi-scale entity cropping with a ReAct-style visual toolchain for cross-modal reasoning; the memory graph for explicit state tracking; and Topology-Aware Policy Optimization (TAPO) that uses graph structure to identify critical paths and prune redundant nodes for better credit assignment.\n\nExperiments on challenging multimodal benchmarks show consistent outperformance over existing baselines on long-horizon reasoning tasks. The work is a single-author submission from Zhongyu Wang, which limits institutional backing; independent reproduction would strengthen the contributions.