TSR Desk · science · 10 October 2026, 01:00 UTC
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
- What
- Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
- Who
- arxiv.org
- When
- 9 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2610.11889
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. It comes from a paper posted to arXiv on 9 October 2026. We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.
Why it counts
Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.