TSR Desk · science · 5 October 2026, 07:00 UTC
$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
- What
- $T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
- Who
- arxiv.org
- When
- 5 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2609.32791
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%. It comes from a paper posted to arXiv on 5 October 2026. Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift.
Why it counts
Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.