ResearchPod Summary
Standard PPO for LLMs typically uses a scalar critic trained via mean-squared-error (MSE) regression to predict expected returns. In reinforcement learning with verifiable rewards (RLVR), where rewards are sparse and binary, even minor miscalibrations in the critic can lead to asymmetric advantage signals that distort policy updates. This paper investigates whether replacing the scalar MSE critic with a categorical predictor—trained as a classification task—can provide a more stable and accurate signal for PPO.
Researchers introduced HL-Gauss PPO, which replaces the final scalar output layer of the critic with a categorical head. This head predicts a distribution over a discretized range of possible values. The critic is trained using a cross-entropy loss against a target distribution constructed via Gaussian smoothing (HL-Gauss). Crucially, the categorical output is decoded back into a scalar expectation before being passed to the standard GAE and PPO update steps. This ensures that the actor-side policy optimization remains unchanged, allowing the researchers to isolate the impact of the critic's training objective.
HL-Gauss PPO consistently outperformed strong PPO and DAPO baselines across mathematical reasoning datasets (e.g., AIME24, AIME25) and tool-augmented search tasks. The authors demonstrate that the categorical critic produces better-calibrated scalar predictions, effectively reducing the extreme asymmetry in advantage signals often seen with MSE critics. Specifically, MSE critics tend to be overconfident on likely-failure prefixes, leading to disproportionately large penalties; the categorical approach mitigates this, resulting in more balanced and lower-variance advantages that facilitate more effective learning.
This work highlights that the critic's learning objective is a critical, yet often overlooked, component of LLM reinforcement learning. By framing value estimation as a classification problem, researchers can achieve significant performance improvements in reasoning tasks without altering the underlying policy optimization algorithm. This provides a simple, drop-in architectural modification that enhances sample efficiency and final model performance in sparse-reward environments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.