TSR Desk · science · 6 October 2026, 01:00 UTC
LEAF: A Living Benchmark for Event-Augmented Forecasting
- What
- LEAF: A Living Benchmark for Event-Augmented Forecasting
- Who
- arxiv.org
- When
- 5 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2605.16358
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
To establish a rigorous evaluation paradigm, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including trend, event, and time series forecasting. It comes from a paper posted to arXiv on 5 October 2026. Large Language Models (LLMs) are increasingly applied to real-world forecasting tasks, yet evaluating their true predictive capability remains compromised by pre-training data contamination and look-ahead leakage in automated search. Existing benchmarks either rely on static contexts, restrict evaluations to narrow environments, or fail to audit auxiliary textual events for future information leakage. LEAF couples a recursive retrieval agent system with dual-agent cross-validation to gather comprehensive, relevant, and temporally aligned auxiliary context. A comprehensive audit across 500 tasks by 47 domain specialists demonstrates that our pipeline suppresses future information leakage from 8.6% to 1.6%. Across extensive evaluations of 16 frontier proprietary and open-weight models on our benchmark, we show that LLMs effectively extract signals from verified events to boost trend and event forecasting.
Why it counts
To establish a rigorous evaluation paradigm, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including trend, event, and time series forecasting.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.