New benchmark confirms AI models still perform poorly at visual perception
| Source: THE DECODER
Tags: PerceptionBench, multimodal, GPT-5.6 Sol, Kimi, visual perception, benchmarks, Moonshot AI
Moonshot AI's PerceptionBench tests visual perception independently of reasoning and knowledge — no frontier model among 16 tested exceeds 60% accuracy, with GPT-5.6 Sol leading at 59.7%, and many assumed reasoning failures traced back to faulty image reading.
Details
Moonshot AI, the team behind the Kimi assistant, has released PerceptionBench, a benchmark that isolates visual perception from logical reasoning and external knowledge. Unlike existing benchmarks that bundle perception, reasoning, and recall together, PerceptionBench breaks vision into 10 atomic skill domains built from real model errors: Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Every task in the benchmark can be answered purely by looking at the image — no reasoning chain or factual recall is required. From an internal pool of 17,000 verified questions, 3,000 tasks are being published: 60% derived from attributed model errors, 40% reformulated using augmented images. The tasks look deceptively simple — identifying where a symbol sits on a clock face, counting objects in a bounded region, distinguishing colors across two items. Across 16 frontier models tested, the highest overall accuracy is 59.7%, scored by GPT-5.6 Sol. Kimi K3 and Claude Fable 5 were also tested. No model crosses 60%. The central conclusion from the authors: a large proportion of failures attributed to reasoning errors in prior studies are actually perception failures that happen at the image-reading stage — before any reasoning step fires.