TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 28 September 2026, 07:00 UTC

DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language

What
DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models
Who
arxiv.org
When
28 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.31103
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. It comes from a paper posted to arXiv on 28 September 2026. Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. Across nine datasets, DepthEvidence achieves the highest average dense $\delta_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.

Why it counts

We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.