Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
An empirical study tested seven agents in a retrieval environment with deliberately failing sources. The agents labelled useless results as useless 97% to 100% of the time, yet most kept querying: the judgment never became a stopping decision. Prompt cues, memory, reasoning modes and budget limits changed when agents stopped, not why. Only a mandatory integration step, forcing an answer after five consecutive useless results, aligned stopping with the evidence. A pre-registered replication on 300 fresh questions confirmed both the disconnect and the fix.
The question the paper asks
The paper "Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions" (arXiv 2610.06191) comes from Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Yaxin Zhou, Ivor Tsang and Bo An. It was submitted on 5 October 2026. Its question is simple. If a tool keeps returning nothing useful, does an agent stop relying on it? Common sense says yes. An agent that can say "this result is useless" should, after several such verdicts in a row, change course or answer with what it has. The authors found that most agents do not.
The team tested seven agents in a retrieval environment. The environment contained deliberately failing sources, so the tool kept returning content with no value. The agents labelled those results useless between 97% and 100% of the time. Their evidence judgments were almost always right. Yet most of them kept querying. The title names the effect: judged useless, queried anyway.
The core finding: judgment and action come apart
The key word is disconnect. Earlier discussions of agent loops often blame the model's inability to read tool output or to assess its quality. This paper rules out that explanation for these agents. A 97% to 100% labelling rate shows that the model knows the result is useless. The failure sits one step later: the judgment does not become a decision to stop.
To measure judgment and action separately, the authors compared stopping behavior after longer runs of results marked useless against stopping after shorter runs. If an agent really used its own judgments to decide, a longer run of useless results should make it more likely to stop. The paper reports that most agents did not show this evidence-sensitive pattern. The design has a practical strength. It does not need an outside label that says when stopping is correct. It tests whether the agent's behavior agrees with the agent's own verdicts.
What failed and what worked
The authors tried the usual remedies: prompt cues, access to memory, different reasoning modes, and budget constraints. The result was sobering. These changes moved the timing of the stop, but not the criterion for stopping. An agent might quit earlier or later, but not because the evidence had piled up against the tool. A budget cap is the clearest case. It cuts the loop by force, with no link to evidence. It is an external brake, not the agent's own decision.
What did work was a mandatory integration step. After five consecutive useless results, the agent must answer from the information it already holds. This moves the stopping decision away from the agent's discretion and into the harness. According to the paper, this aligned stopping behavior with the evidence. The authors then ran a pre-registered replication on 300 fresh questions. It confirmed both parts of the story: the disconnect is real, and the integration step fixes it. Pre-registration fixes the hypotheses and analysis plan before the data is seen, which lowers the risk of selective reporting.
A plausible reading of the mechanism
The abstract does not give a full mechanistic account. The following is our inference from the reported results, not the authors' claim. One possibility is that training data contains many trajectories where one more query was harmless, so calling the tool again becomes the default action.
A written verdict such as "useless" is only text. It need not shift the probability of the next action. A second possibility is that judging is local to a single result, while stopping needs a running tally across the whole trajectory, which is harder for small models. The integration step may work because it hands the tally to an external counter and asks the model only to answer.
What this means for builders
First, cost. Every needless tool call adds latency and spend. In retrieval-augmented generation, web browsing and code search, a wasted loop is a real bill. Second, reliability. An agent that spins on a dead source can time out, or it can produce a thin answer at the last moment. Third, design. The findings suggest that teams should not count on the model to work out when to stop. A deterministic rule in the orchestration layer, such as "after N useless results, force an answer", is easy to build, works in any agent framework, and needs no retraining.
The work also bears on evaluation. Many agent benchmarks score only the final answer and rarely measure wasted tool calls. This paper offers a reusable diagnostic: first check whether the judgments are accurate, then check whether behavior follows them. Separating the two shows whether a failure comes from misreading or from not acting.
Limits and open questions
Several limits deserve plain mention. The abstract says the study used models in the 7 to 8 billion parameter range and does not name them, so readers should consult the full text before assuming the result carries over to larger frontier models. The test environment uses sources that fail on purpose. Real failures are murkier: results can be partly useful, and a source can be good on Monday and bad on Tuesday.
A fixed threshold of five may fire too early or too late there. Forcing an answer also means answering with thin evidence, and the cost to answer quality needs task-level study. Finally, the finding that prompts and budgets change timing but not criteria does not prove that nothing better exists. Training that directly rewards stopping because of evidence, for example with reinforcement learning, is a natural next step.
Takeaway
The paper takes a vague complaint, "the agent got stuck in a loop", and splits it into two parts. The model can judge.
It does not act on its own judgment. The fix the authors propose is simple and reproducible. The lesson for engineers is to treat the stopping decision as a first-class part of agent design, and not to leave it to the model's good sense.