Thought without systematicity? Evaluating reasoning models on rule induction tasks

| Source: arXiv AI

Tags: reasoning models, systematicity, cognitive science, benchmarking, rule induction, evaluation

A study by Schug and Lake (NYU) finds that current reasoning models consistently fail on structurally equivalent variants of tasks they can solve—exposing that model capabilities may be tightly context-bound rather than reflecting genuine systematicity, a fundamental property of human cognition.

Details

Systematicity is foundational to human reasoning: understanding one concept implies understanding structurally related variants. This paper asks whether current reasoning models exhibit the same property using established rule induction tasks from cognitive science.\n\nSchug and Lake create structurally equivalent task variants through isomorphisms like recombination and substitution—tasks that require the same underlying cognitive operation applied to different surface forms. If models were truly systematic, performance should be consistent across these variants.\n\nThe findings are direct: despite being able to correctly solve a task, models frequently fail on structurally equivalent variants. The authors argue this indicates that observed capabilities are tightly bound to specific evaluation contexts rather than reflecting robust underlying reasoning. This has immediate implications for practitioners: high benchmark scores may not predict performance on surface reformulations of the same problem in real deployments. Code is available.