From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
| Source: arXiv AI
Tags: video generation, VWG-Bench, Vid-PRE, video reasoning, ECCV 2026, benchmark
VWG-Bench exposes a stark gap in video generation: models score well on visual quality but consistently fail logic-heavy and rule-constrained tasks across 9 reasoning dimensions. Vid-PRE, a RL-trained prompt rewriter, closes much of this gap without touching model weights.
Details
Video generation models produce compelling footage, but do they actually reason about the world? VWG-Bench (Video World Generalist Benchmark) stress-tests this across 38 fine-grained tasks in 9 reasoning dimensions — symbolic rules, physical laws, intentional goals, and more. The evaluation uses a three-level VLM-as-Judge protocol that independently scores video fluency, rule adherence, and goal realization, separating perceptual quality from cognitive correctness. Results on leading models reveal the gap: strong rendering scores, consistent failure on logic-heavy tasks. To address this, the team introduces Vid-PRE (Video Prompt Reasoner and Enhancer) — a model-agnostic prompt rewriter trained via reinforcement learning using only text-based rewards. It rewrites the user's prompt into a concise, constraint-aware version before passing it to any video generator, offloading the cognitive burden to a dedicated VLM. No architectural changes to the generator are needed. Accepted at ECCV 2026 with data and code publicly available. The benchmark and rewriter together provide both a diagnosis tool and a practical fix for video generation reasoning failures — useful for teams shipping video generation products.