TSR Desk · science · 17 September 2026, 01:00 UTC
Post-Reasoning: Improving the Performance of Non-Thinking Models at No Cost
- What
- Post-Reasoning: Improving the Performance of Non-Thinking Models at No Cost
- Who
- arxiv.org
- When
- 16 September 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2605.06165
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
In this work, we propose Post-Reasoning, a simple yet effective approach that improves instruction-tuned models by conditioning them to justify their answers after generating the final response. It comes from a paper posted to arXiv on 16 September 2026. As the widespread adoption of Large Language Models (LLMs) accelerates, token consumption from intermediate reasoning traces increasingly contributes to inference latency and operational cost. Recent studies suggest that many real-world tasks require little to no explicit reasoning, with additional reasoning sometimes even degrading performance. By design, it enables the final answer to be obtained without additional latency or token cost, while still improving performance through simple instruction augmentation. We evaluate Post-Reasoning across 117 model--benchmark settings spanning 13 open and proprietary models, 4 model families, and 9 diverse reasoning and knowledge-intensive benchmarks, including AMC, HMMT, GSM8K, GPQA, MMLU-Pro, and BIG-Bench Hard. Post-Reasoning improves performance in over 88.19% of evaluated settings, achieving a mean relative improvements of 17.37%. Furthermore, we propose supervised post-reason tuning, which further improves performance in over 91.11% of evaluated settings, and exceeds the prompt-based post-reasoning baseline by an average of 8.01%, demonstrating that post-reasoning can be effectively internalized through training. Ultimately, Post-Reasoning establishes a new performance ceiling for direct-answer capabilities.
Why it counts
In this work, we propose Post-Reasoning, a simple yet effective approach that improves instruction-tuned models by conditioning them to justify their answers after generating the final response. Post-Reasoning improves performance in over 88.19% of evaluated settings, achieving a mean relative improvements of 17.37%. Furthermore, we propose supervised post-reason tuning, which further improves performance in over 91.11% of evaluated settings, and exceeds the prompt-based post-reasoning baseline by an average of 8.01%, demonstrating that post-reasoning can be effectively internalized through training.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.