TSR Desk · science · 14 September 2026, 07:00 UTC
Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
- What
- Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
- Who
- arxiv.org
- When
- 14 September 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2609.12446
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. It comes from a paper posted to arXiv on 14 September 2026. Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.
Why it counts
We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.