Karn Tiwari, Varnith Chordia, Prathosh A.P.
5 min
Abstract
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines \emph{which trajectories} receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by $+1.7$ and $+1.8$ points and pass@8 by $+1.6$ and $+5.7$ points, respectively. On mathematics, avg@8 remains within $0.5$ points of GRPO while pass@8 improves by $+1.1$ and $+3.9$ points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.
Alex: Then where would a careful referee push? The verifier is binary, and the whole gate hangs on it.
Sam: The verifier is the bottleneck. A false negative silences the teacher on a trajectory that may have been correct. The group helps somewhat, because the solve rate is computed over several samples and so acts as a smoother rather than resting on one pass-fail judgment. But it doesn't remove the dependence.
Alex: And the teacher? It's a frozen model. If it has a systematic bias, say a preference for an inefficient reasoning style, does the student just inherit it?
Sam: Yes, that's a known trade-off. The student is pulled toward the teacher's distribution on the failed subset, so bias propagates. The tanh cap limits how hard any token-level disagreement can pull, and the gate means the student only aligns when it's clearly failing. But it's a soft constraint, not protection. The goal is better reasoning coverage, not discovering strategies the teacher doesn't already have.
Alex: And because the comparison is at every token, the student gets the whole reasoning chain, not just the final answer.
Sam: Yes. It's the difference between being told the answer was wrong and being shown which step in the derivation went off the rails. That dense signal is what outcome-only reinforcement learning lacks on hard prompts.
Alex: So the takeaway is narrow but sensible. The gate is a fallback for when environment feedback is too sparse to learn from.
Sam: That's how I'd put it. It concentrates local supervision where outcome feedback is weakest, with the caveats that it depends on the verifier and the teacher. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.