TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 10 September 2026, 01:00 UTC

Do Web Agents Investigate Before They Decide?

What
Do Web Agents Investigate Before They Decide?
Who
arxiv.org
When
9 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2602.05354
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Second, procedural hints improve investigation but do not consistently improve decisions on Wikipedia tasks, where decisive evidence often contradicts surface impressions. It comes from a paper posted to arXiv on 9 September 2026. Autonomous web agents are increasingly deployed in moderation and policy enforcement, where correct decisions often depend on evidence that is not immediately visible and must be actively investigated. Yet existing benchmarks largely assume task critical information is immediately accessible. They do not measure investigative competence: recognizing when visible context is insufficient, retrieving hidden evidence, and integrating it into a final decision. We introduce MIRAGE, a benchmark of 750 multi step decision tasks across three domains: Wikipedia Forensics, Shopping Admin adjudication, and Reddit Moderation. Each task has two layers: a visible surface context that often points to the wrong action, and a hidden context, reachable only by active investigation, that contains the decisive evidence. We decompose agent performance into Investigation, Reasoning, and Decision Accuracy, complemented by an Investigative Hallucination Rate. We evaluate eight LLM agents across two model generations. Three patterns emerge. First, agents reach relevant pages but rarely extract the decisive evidence on them. Third, 12.6% of trajectories cite fabricated facts. We call these the Navigation Discovery Gap, Collapse under Contradiction, and Investigative Hallucination. These patterns persist across model scale, generation, and reasoning architecture.

Why it counts

Second, procedural hints improve investigation but do not consistently improve decisions on Wikipedia tasks, where decisive evidence often contradicts surface impressions. First, agents reach relevant pages but rarely extract the decisive evidence on them.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.