The Maivia Gazette

Verified AI news, every morning

Research

GPT-6 Astra triples Fable 5.1's vending profits and is first to beat the human baseline on all five Drone-Bench tasks

Andon Labs reports a $15,515 average bank balance on Vending-Bench and drone code that finds and follows a specific person, though success rates remain unreliable.

A red vending machine sits on a sunlit rooftop while a small drone hovers above it.
AI-generated illustration, not event photography.

Andon Labs has tested OpenAI's GPT-6 Astra on two agent benchmarks that measure very different capabilities, and reports that the model outperforms all previous frontier models on both. Vending-Bench simulates running a vending machine business over a long period, including buying inventory and managing money. Astra averaged a final bank balance of $15,515, nearly three times the result for Anthropic's Claude Fable 5.1. Drone-Bench measures whether a model can write software for a physical system. Andon Labs says Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks, including writing code that lets a drone autonomously find and follow a specific person. The lab cautions that Astra's success rate on the drone tasks remains unreliable, so the headline result reflects best attempts rather than consistent performance. The findings add to a run of early evaluations suggesting Astra handles long-horizon autonomous tasks and physical-world reasoning better than its predecessors. They also underline a dual-use concern: the same capability that lets a model program a drone to track a person is directly applicable to surveillance. As with other third-party benchmarks published shortly after a model's release, the results come from a single lab and have not yet been independently replicated.

Sources

  1. The DecoderGPT-6 Astra pilots a surveillance drone and runs a business on its ownPublished · fetched

Also in this edition