The Maivia Gazette

Verified AI news, every morning

Research

Paired study finds hosted decision model Jev beats open-weight Laya on most agent choices, and neither can route models

In tests on 11 agent decision points, Laya changed 30 percent of its answers when the order of options was reversed.

A row of iron switch levers in a signal box overlooking a dusky rail yard where the tracks split into many paths.
AI-generated illustration, not event photography. The motion is AI-generated from the still.

A new arXiv paper by Jiawei Li evaluates so-called System-1 decision models. These are small models that answer typed questions inside agent workflows in a single forward pass and return class probabilities. Examples include which model to call, which tool to use, whether retrieved text is relevant, and whether an input contains a prompt injection. The study compares an open-weight model, Laya, with a hosted model, Jev, on 11 agent decision points built from 18 public sources. It uses 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests and reproducibility checks across hardware and days. Jev was significantly more accurate on 9 of the 11 decision points, by margins of 10.8 to 46.0 percentage points. Neither model beat chance on zero-shot model routing, and the two tied on gating retrieved text for relevance. Laya changed 30 percent of its answers when the order of the options was reversed, and its accuracy fell sharply when it faced many similar candidates: 31 percent with 50 nearest-neighbour tools. Jev scored 98 percent on items with a single correct tool. The author also audited the study's own analysis pipeline and reports finding three errors in it. The results suggest that developers should test such models for robustness before relying on them as cheap steps in production agent workflows.

Sources

  1. arXiv.orgFast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent HarnessesPublished · fetched

Also in this edition