TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 5 October 2026, 07:00 UTC

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under

What
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
Who
arxiv.org
When
5 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2607.02814
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

We introduce SovereignNegotiation-Bench, a controlled benchmark that operationalizes five such duties from agency law--loyalty, obedience to actual authority, confidentiality, candor and diligence--as deterministic checks on episode logs; the first three enter a single headline metric. It comes from a paper posted to arXiv on 5 October 2026. Personal AI agents are beginning to negotiate for people, from refunds and bills to deposits and sales. A human agent in that position is judged by the duties owed to the principal, not by whether a deal was struck. The benchmark contains 1,764 paired scenarios (252 situations in 18 consumer and peer-to-peer domains, each under 7 counterparty tactics). The counterparty's economics are a fixed function of the agent's structured actions and of the disclosures detected in its messages, so outcomes are comparable across agents and a disclosed limit has a measurable, causal price. A simulated principal grants or withholds consent and tightens its mandate midepisode. Rule-based agents show that the benchmark is solvable from the observable state (92% faithful success) and that a single disclosing sentence erases the entire negotiated surplus (0.70 to 0.00). Across 17 open-weight models, faithful success ranges from 6% to 75%; models disclose the principal's reservation value in 2-80% of episodes, agree or share a protected document without a required approval in 2-23%, and follow an instruction injected into the counterparty's message in 5-57% of injection episodes. Deal rate ranks models much like faithful success does, but it does not certify individual agreements: pooled over models, 48% of the agreements breach at least one duty (18-96% per model). Within families, faithful success tends to rise with size, but no size trend is significant, and on the model we test, neither prompting nor a code-level guard raises faithful success substantially. Code, scenarios and all episode logs will be released.

Why it counts

We introduce SovereignNegotiation-Bench, a controlled benchmark that operationalizes five such duties from agency law--loyalty, obedience to actual authority, confidentiality, candor and diligence--as deterministic checks on episode logs; the first three enter a single headline metric.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.