GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks

| Source: THE DECODER

Tags: GPT-6 Astra, OpenAI, robotics, spatial reasoning, StationeryBench, MolmoAct2

GPT-6 Astra completed 7 out of 100 physical robot tasks on the new StationeryBench benchmark using dual-arm robots — with a median progress score of 46/100 — while competitor MolmoAct2 completed zero; Cornell researcher Yoav Artzi called it a 'step change in spatial reasoning.'

Details

StationeryBench is a new robotics benchmark pitting AI models against five desk-object manipulation tasks: uncapping a marker, pouring out paper clips, passing a ruler between two robot arms, and similar tasks. Both models controlled the same dual-arm YAM robots across 200 trials (100 per model).\n\nGPT-6 Astra fully completed 7 tasks; Ai2's MolmoAct2 completed zero. More telling is the median progress score: Astra reached 46/100, MolmoAct2 only 12/100. All results, videos, and code are published on GitHub.\n\nYoav Artzi (Cornell / Google DeepMind) described Astra as a 'step change in spatial reasoning.' On the still-unpublished REMAP benchmark, GPT-6 Astra approaches human-level accuracy, though Artzi notes it still falls short in other scenarios. He suspects OpenAI trained on large volumes of 3D data such as Blender scenes, consistent with Astra's particular strength on 3D tasks.\n\nThe benchmark results are notable context for OpenAI's long-term consumer robotics ambitions. However, absolute completion rates (7%) are still low, and the task set is narrow — desk objects in controlled conditions rather than open-world environments.