GPT-6 Astra completes 7 of 100 desk-manipulation tasks where Ai2's MolmoAct2 finishes none
On the new StationeryBench, Astra's median progress score is 46 to MolmoAct2's 12, and Cornell's Yoav Artzi calls the result a 'step change in spatial reasoning.'

Early benchmarks suggest GPT-6 Astra is markedly better at spatial reasoning than prior systems. A new robotics benchmark called StationeryBench compares OpenAI's model against Ai2's MolmoAct2 on five desk-object tasks such as uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials. Astra fully completed 7 of 100 tasks while MolmoAct2 completed zero, and Astra's median progress score reached 46 out of 100 against 12 for MolmoAct2. All results, videos, and code are published on GitHub. Yoav Artzi, an AI researcher at Cornell and Google DeepMind, describes Astra as a 'step change in spatial reasoning.' On the still-unpublished REMAP benchmark, he says the model reaches accuracy close to human level, while cautioning that 'even ASTRA doesn't get to what humans do in other scenarios.' Artzi suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes, which would fit its particular gains on 3D tasks. The absolute completion rate remains low, so the results indicate direction rather than readiness for real work. They matter because OpenAI has stated long-term plans to build its own consumer robots, and spatial competence in a general model is a prerequisite for that ambition.