TSR Desk · science · 3 October 2026, 01:00 UTC
ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents
- What
- ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents
- Who
- arxiv.org
- When
- 2 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2606.00341
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
We introduce ROGUE, a benchmark in which agents are asked to complete realistic computer-use tasks but encounter controlled conflicts with human control, shutdown, or explicit resource restrictions. It comes from a paper posted to arXiv on 2 October 2026. As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although much work has focused on agent safety in the presence of an adversary, we study corrigibility: whether agents remain amenable to human correction, interruption, or shutdown while pursuing benign tasks. We then evaluate whether agents violate these constraints in pursuit of task completion: overriding the human, accessing restricted passwords, or rewiring shutdown. We find that most frontier models tested frequently bypass user interruptions or restrictions under the evaluated conditions, and that text-only evaluations can underestimate failures during agentic execution. Further, independent task capability does not by itself imply greater corrigibility. Finally, even when a parent agent behaves corrigibly, safety constraints may fail to propagate to the subagents it creates.
Why it counts
We introduce ROGUE, a benchmark in which agents are asked to complete realistic computer-use tasks but encounter controlled conflicts with human control, shutdown, or explicit resource restrictions.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.