TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 10 October 2026, 01:00 UTC

Evaluating Exact Output and Checkpoint-State Prediction in Real Programs

What
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
Who
arxiv.org
When
9 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2610.11889
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. It comes from a paper posted to arXiv on 9 October 2026. We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.

Why it counts

Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.