TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 15 September 2026, 01:00 UTC

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

What
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
Who
arxiv.org
When
14 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.12808
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

We introduce K-Bench, a benchmark that scores LLM unlearning under agentic deployment. It comes from a paper posted to arXiv on 14 September 2026. Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent. K-Bench inspects all six channels a ReAct agent exposes, including its chain-of-thought (CoT), tool calls and tool observations, and elicited summary. A query counts as leaked if the secret appears in any of them. Each experiment places the secret in exactly one of the agent's three sources (the weights, the prompt, or the retrieval store). The K-Score is computed separately for each source and credits forgetting only when the agent remains usable. Clearing the answer channel does not make the secret unrecoverable. On structured retrieval, the secret stays verbatim in the tool-observation channel and the aggregate leak rate is unchanged. When the secret lives in the prompt or the retrieval store, TOFU and MUSE report no leakage, while the deployed agent still leaks it on 22--86\% of queries. When the secret is in the weights, none of the twenty evaluated published methods demonstrably removes it, and only an input-corruption intervention reaches selective forgetting under the evaluated observer. The top-ranked method changes across base models. A refusal-tuning method resists the evaluated extraction without verified knowledge removal.

Why it counts

We introduce K-Bench, a benchmark that scores LLM unlearning under agentic deployment.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.