Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation
| Source: arXiv AI
Tags: video generation, benchmark, multimodal, CutCraft, video editing, evaluation
CutCraft is the first benchmark testing whether AI video generators actually execute professional editing techniques — J-cuts, L-cuts, shot transitions — finding that current SOTA models produce plausible videos but consistently fail to follow editorial instructions.
Details
Multi-shot video generation has improved in coherence, but coherence is not the same as editorial competence. CutCraft, introduced by researchers from Alibaba and collaborating institutions, tests whether models can execute specific editing techniques: shot structure, transition grammar, audio-video cut timing (J-cuts, L-cuts), and montage structure. The benchmark pairs structured multi-shot prompts with explicit editing specifications, then evaluates them with a hierarchical hybrid framework combining shot-structure alignment, expert-model metrics, tool-grounded multimodal judgment, and rubric-based QA. An agentic baseline decomposes generation into planning, shot-level synthesis, and post-hoc composition to realize editing semantics explicitly. Across 13 closed- and open-source SOTA models, a consistent gap emerges: systems generate visually plausible multi-shot video but fail to reliably execute editorial instructions. Specific failure modes include unstable shot structures, weak transition control, and sharp degradation on higher-order montage. Aesthetic quality correlates only weakly with editing compliance. For teams building AI video production tools, CutCraft offers a concrete evaluation layer that prior benchmarks missed. The benchmark and metrics are publicly available.