ResearchPod Summary
As large language models increasingly rely on step-by-step reasoning to solve complex problems, they often generate redundant or overly verbose reasoning traces. Existing On-policy Distillation (OPD) methods improve reasoning accuracy but typically rely on single-distribution matching, which fails to explicitly model preferences between different reasoning modes. This paper investigates whether a contrastive approach can guide models toward more efficient, concise reasoning without sacrificing performance.
The authors introduce Contrastive On-Policy Distillation (COPD). Instead of imitating a single teacher distribution, COPD uses a frozen teacher to evaluate the student's generated tokens under two distinct contexts: a 'light-thinking' instruction (encouraging brevity) and a 'heavy-thinking' instruction (encouraging thoroughness). The difference in log-probabilities between these two evaluations serves as a token-level advantage signal. This signal is then used to update the student policy using a clipped PPO-style objective, effectively pushing the model to favor tokens that align with the light-thinking preference.
COPD consistently improves the accuracy-to-length trade-off across nine multimodal benchmarks. By shifting the student toward more efficient reasoning, the model achieves significant reductions in average response length—often by more than 50%—while simultaneously improving or maintaining accuracy. The authors also demonstrate that this contrastive mechanism can be applied as a self-distillation (COPSD) technique, where the model uses its own previous snapshots to generate the contrastive signal, proving that the compression benefit is derived from the relative preference mechanism rather than just the scale of the teacher model.
This research provides a scalable, label-free method for optimizing the inference cost of reasoning models. By transforming teacher scoring into a relative preference task, it allows developers to compress reasoning traces adaptively. This is particularly valuable for deploying reasoning models in resource-constrained environments where reducing token consumption is critical for latency and cost reduction.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.