Paired study finds hosted decision model Jev beats open-weight Laya on most agent choices, and neither can route models
In tests on 11 agent decision points, Laya changed 30 percent of its answers when the order of options was reversed.

A new arXiv paper by Jiawei Li evaluates so-called System-1 decision models. These are small models that answer typed questions inside agent workflows in a single forward pass and return class probabilities. Examples include which model to call, which tool to use, whether retrieved text is relevant, and whether an input contains a prompt injection. The study compares an open-weight model, Laya, with a hosted model, Jev, on 11 agent decision points built from 18 public sources. It uses 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests and reproducibility checks across hardware and days. Jev was significantly more accurate on 9 of the 11 decision points, by margins of 10.8 to 46.0 percentage points. Neither model beat chance on zero-shot model routing, and the two tied on gating retrieved text for relevance. Laya changed 30 percent of its answers when the order of the options was reversed, and its accuracy fell sharply when it faced many similar candidates: 31 percent with 50 nearest-neighbour tools. Jev scored 98 percent on items with a single correct tool. The author also audited the study's own analysis pipeline and reports finding three errors in it. The results suggest that developers should test such models for robustness before relying on them as cheap steps in production agent workflows.