TSR Desk · physics · 13 September 2026, 01:00 UTC
When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series
- What
- When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting
- Who
- arxiv.org
- When
- 12 September 2026, 04:00 UTC
- Category
- Physics
- Primary source
- https://arxiv.org/abs/2609.11282
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
This is an information-theoretic question, but to evaluate whether information-theoretic metrics can reliably measure the predictive value an annotation provides, a ground truth benchmark is needed, and none currently exist. It comes from a paper posted to arXiv on 12 September 2026. Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? We create a synthetic time series signal with annotations in three categories: semantically correct, incorrect, and irrelevant. Because the data generation process is fully controlled, ground-truth information content is known exactly, enabling principled evaluation of six complementary mutual information estimators (KSG, MINE, InfoNCE, CCA, PID and V-information). We show that all six estimators identify correct annotations as most informative, and are able to audit the quality of mixed text corpora, choosing the annotations that result in the best downstream forecasting results without the need for model training. Our benchmark identifies limitations of each estimator, and these are validated on seven real-world datasets, which show how estimator performance differs on weak signals. Finally, we establish practical rules for implementing these metrics for annotation auditing and fusion selection.
Why it counts
This is an information-theoretic question, but to evaluate whether information-theoretic metrics can reliably measure the predictive value an annotation provides, a ground truth benchmark is needed, and none currently exist. Our benchmark identifies limitations of each estimator, and these are validated on seven real-world datasets, which show how estimator performance differs on weak signals.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.