TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 10 October 2026, 01:00 UTC

Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark

What
Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
Who
arxiv.org
When
9 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2610.10740
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage. It comes from a paper posted to arXiv on 9 October 2026. Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning.

Why it counts

The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.

Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark · The Singularity Report