TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 29 September 2026, 01:00 UTC

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level

What
UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
Who
arxiv.org
When
28 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.30928
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. It comes from a paper posted to arXiv on 28 September 2026. Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.

Why it counts

Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.