TSR Desk · science · 8 October 2026, 01:00 UTC
COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
- What
- COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
- Who
- arxiv.org
- When
- 7 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2610.04378
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Combined input and context scanning reduces mean attack success by 49.6% on an eight-attack defense subset relative to the undefended setting. It comes from a paper posted to arXiv on 7 October 2026. Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with guardrails, orchestration, and general task capability. We introduce COPEX (COntext Provider EXploitation), a controlled benchmark that isolates the model as an MCP client by fixing the surrounding agent stack and varying only the tool-selecting model. COPEX covers 25 attack types instantiated as 125 scenarios across four entry surfaces: model/agent, client, server/tool, and transport. Across nine models and 3,375 trials, the mean attack success rate is 64.4%, with surface-level means ranging from 58.3% to 71.4%. Some client- and transport-level attacks succeed partly outside the model's observation or control, separating system exposure from model susceptibility. The benchmark is available at https://github.com/inspire-center/copex.
Why it counts
Combined input and context scanning reduces mean attack success by 49.6% on an eight-attack defense subset relative to the undefended setting.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.