TSR Desk · science · 17 September 2026, 01:00 UTC
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
- What
- $\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
- Who
- arxiv.org
- When
- 16 September 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2609.13602
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
We introduce $\tau$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments. It comes from a paper posted to arXiv on 16 September 2026. Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. A matched text agent passes all tasks, but four voice configurations achieve robust exact success from 0.14 to 0.41. Agents increase verification for hard and unfamiliar entities and sometimes for incorrect captures, but not for their weakest caller voice; only 24 to 37 percent of verified errors are repaired. A scaffold that prescribes spelling, read-back, correction, and confirmation raises robust Pass$^3$ by 14 to 31 points, at a cost of 21 to 28 seconds per call. Realisms such as spelling variations and restarts do not detectably affect exact success; mispronunciation increases repair effort. These results identify strategy selection and successful recovery as the central bottlenecks in exact spoken entity collection.
Why it counts
We introduce $\tau$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.