MindTopo: Can Foundation Models Reason in Topological Space?

| Source: arXiv AI

Tags: benchmark, spatial reasoning, topology, multimodal LLMs, MindTopo, Qwen3-VL, planning

MindTopo — 11,030 instances across continuity, separation, order, enclosure, and knot tasks — finds all 14 MLLMs tested fall far below human performance on topological planning, with fine-tuning improving reasoning more than planning and video generation failing to preserve topology across transitions.

Details

Spatial reasoning in AI evaluations has focused heavily on metric properties (distance, angle, shape) while topological relations — invariant under continuous deformation, like whether objects are enclosed or connected — have been largely ignored. MindTopo fills this gap with the first benchmark specifically targeting topological intuition.\n\nFive properties grounded in cognitive science and formal topology are covered: continuity, separation, order, enclosure, and knots. Each is tested at two cognitive levels: reasoning (identify or infer topological relations) and planning (closed-loop agent selecting environment actions). With 11,030 instances across 13 procedurally generated task types, the benchmark supports systematic comparison at varying difficulty levels.\n\nResults across 14 MLLMs are consistent: all perform significantly better on reasoning than planning, and the best model remains far below observed human performance. Supervised fine-tuning and RL on Qwen3-VL-2B improve reasoning more than planning — suggesting planning failures are not simply about lacking topological knowledge.\n\nVideo generation augmentation (3 models tested) retains local cues and reaches plausible endpoints, but audited rollouts do not reliably preserve topological invariants across action sequences. This suggests the planning gap is structural, not solvable by adding visual generation alone.