TSR Desk · science · 7 October 2026, 01:00 UTC
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward
- What
- Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
- Who
- arxiv.org
- When
- 6 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2608.16072
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. It comes from a paper posted to arXiv on 6 October 2026. Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. Consequently, an objective whose rewards are already near their upper bound can retain substantial influence when its rewards still vary within rollout groups. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This changes the relative contribution of each objective according to its observed reward headroom. We further derive an exact condition under which saturation aware reweighting reverses the sign of a rollout's aggregate advantage. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels. Additional experiments characterize the accompanying reward tradeoffs and sensitivity to corrupted rewards.
Why it counts
Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.