The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

| Source: arXiv AI

Tags: multimodal, benchmark, GPT-4o, Gemini, CVPR, cross-modal-reasoning

A new multimodal benchmark asks models to infer words from pen-scratch audio and hand-movement video alone — humans score over 80%, while GPT-4o and Gemini 2.5 Pro both fail to surpass 10%, revealing a fundamental gap in cross-modal causal reasoning.

Details

The Unwritten Benchmark defines a novel task called acousto-kinematic word inference: given only the audio of pen scratches and video of a hand moving across paper — with no visible ink — models must identify the word being written across three different writing styles. It targets the ability to synthesize complementary perceptual cues dynamically, a capability humans apply without effort. Human participants achieve over 80% ordered letter accuracy. GPT-4o and Gemini 2.5 Pro both score below 10%. More striking: providing both modalities together often degrades model performance compared to either alone. The authors call this a 'paradoxical fusion effect' and attribute it to a fundamental breakdown in cross-modal causal reasoning — these models cannot figure out that audio and video from the same pen movement should reinforce rather than conflict. The benchmark is accepted to CVPR Findings 2026. The roughly 70-percentage-point gap between human and machine performance on a task most humans find trivial suggests current multimodal architectures are not assembling perceptual evidence the way humans do — they classify patterns in each modality separately rather than inferring a common causal source.