ResearchPod Summary
Large language models often generate excessively long reasoning traces, incurring high computational costs. While researchers have attempted to mitigate this by adding length-penalizing rewards to the Group Relative Policy Optimization (GRPO) algorithm, these methods frequently suffer from reward collapse—a failure mode where the model suddenly produces extremely short, low-quality outputs. This paper investigates the mechanistic root of this instability to enable stable efficiency training.
Through a systematic evaluation of various reward configurations, the authors identify two distinct pathways for reward collapse:
To address these issues, the authors introduce Adaptive Correct-Only Efficiency Reward (ACOER). ACOER employs three synergistic mechanisms:
Empirical results on benchmarks like MATH-500 show that ACOER reduces token generation by over 60% while maintaining accuracy comparable to the base model, providing a robust framework for efficient reasoning.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.