TSR Desk · science · 8 October 2026, 01:00 UTC
Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in
- What
- Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
- Who
- arxiv.org
- When
- 7 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2610.05526
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. It comes from a paper posted to arXiv on 7 October 2026. Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.
Why it counts
SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.