TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 8 October 2026, 01:00 UTC

Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in

What
Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
Who
arxiv.org
When
7 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2610.05526
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. It comes from a paper posted to arXiv on 7 October 2026. Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.

Why it counts

SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.