The Maivia Gazette

Verified AI news, every morning

Research

GPT-6 Astra completes 7 of 100 desk-manipulation tasks where Ai2's MolmoAct2 finishes none

On the new StationeryBench, Astra's median progress score is 46 to MolmoAct2's 12, and Cornell's Yoav Artzi calls the result a 'step change in spatial reasoning.'

Two robotic grippers hand a wooden ruler across a white desk as paper clips spill from a jar.
AI-generated illustration, not event photography.

Early benchmarks suggest GPT-6 Astra is markedly better at spatial reasoning than prior systems. A new robotics benchmark called StationeryBench compares OpenAI's model against Ai2's MolmoAct2 on five desk-object tasks such as uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials. Astra fully completed 7 of 100 tasks while MolmoAct2 completed zero, and Astra's median progress score reached 46 out of 100 against 12 for MolmoAct2. All results, videos, and code are published on GitHub. Yoav Artzi, an AI researcher at Cornell and Google DeepMind, describes Astra as a 'step change in spatial reasoning.' On the still-unpublished REMAP benchmark, he says the model reaches accuracy close to human level, while cautioning that 'even ASTRA doesn't get to what humans do in other scenarios.' Artzi suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes, which would fit its particular gains on 3D tasks. The absolute completion rate remains low, so the results indicate direction rather than readiness for real work. They matter because OpenAI has stated long-term plans to build its own consumer robots, and spatial competence in a general model is a prerequisite for that ambition.

Sources

  1. The DecoderGPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarksPublished · fetched

Also in this edition