On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines \emph{which trajectories} receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by $+1.7$ and $+1.8$ points and pass@8 by $+1.6$ and $+5.7$ points, respectively. On mathematics, avg@8 remains within $0.5$ points of GRPO while pass@8 improves by $+1.1$ and $+3.9$ points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.
Sam: A paper called DiffGate, by Karn Tiwari and colleagues, sends teacher guidance only to the trajectories that fail the verifier, and scales it by how hard the prompt is. In their experiments that improved pass-at-eight across every tested setting in math and code.
Alex: Before the mechanism, what's the gap? Why isn't GRPO enough on its own?
Sam: GRPO has a blind spot. When every sample in a group fails, the group-relative advantage vanishes, so the model gets no learning signal on exactly the prompts where it most needs one. Standard on-policy distillation has the opposite problem. It gives dense token-level feedback, but it's outcome-agnostic, so it can push the student away from a perfectly valid reasoning path just because it differs from the teacher's style.
Alex: So one signal is sparse and coarse, the other dense but potentially misguided. How do you combine them without the two conflicting?
Sam: The verifier acts as a gate. Teacher guidance applies only to trajectories that fail the check, and if the student solves the problem, the teacher stays silent. The strength of that guidance is scaled by the group solve rate, so the teacher intervenes hardest when the student is struggling on a difficult prompt. The guidance itself is a token-level log-probability gap between teacher and student, passed through a bounded tanh so no single token can dominate the update.
Alex: Is the bounding doing real work, or is it a convenience? Raw log-ratio signals can be badly behaved.
Sam: The authors say their diagnostics show that raw teacher signals produce gradient spikes and policy entropy collapse, and that bounding keeps updates in a stable range. That's the stated reason for the design. I'd read it as a diagnostic argument rather than a full ablation of alternatives, at least from what's described here.
Alex: And the headline numbers? Is this a marginal gain or something larger?
Sam: Pass-at-eight improves in all tested settings against a matched GRPO baseline, with a notable increase in solution coverage on code. I'd stress that this is a coverage metric. It tells you the model can find a correct solution across several attempts, not that the average attempt is better.
Alex: That's a useful distinction. Coverage is what you'd expect this method to move, since it targets prompts where the model currently solves nothing.
Sam: Right, and that also explains why the gain is larger on code than on math. On code, a whole batch can fail the unit tests, the group-relative advantage is zero, and GRPO has nothing to learn from. That's where the teacher fills in. Math benchmarks have higher solve rates, so the verifier already provides frequent updates and the gate matters less. That's the authors' explanation, and it's plausible, but it's a mechanism story rather than something they isolate.
Alex: Then where would a careful referee push? The verifier is binary, and the whole gate hangs on it.
Sam: The verifier is the bottleneck. A false negative silences the teacher on a trajectory that may have been correct. The group helps somewhat, because the solve rate is computed over several samples and so acts as a smoother rather than resting on one pass-fail judgment. But it doesn't remove the dependence.
Alex: And the teacher? It's a frozen model. If it has a systematic bias, say a preference for an inefficient reasoning style, does the student just inherit it?
Sam: Yes, that's a known trade-off. The student is pulled toward the teacher's distribution on the failed subset, so bias propagates. The tanh cap limits how hard any token-level disagreement can pull, and the gate means the student only aligns when it's clearly failing. But it's a soft constraint, not protection. The goal is better reasoning coverage, not discovering strategies the teacher doesn't already have.
Alex: And because the comparison is at every token, the student gets the whole reasoning chain, not just the final answer.
Sam: Yes. It's the difference between being told the answer was wrong and being shown which step in the derivation went off the rails. That dense signal is what outcome-only reinforcement learning lacks on hard prompts.
Alex: So the takeaway is narrow but sensible. The gate is a fallback for when environment feedback is too sparse to learn from.
Sam: That's how I'd put it. It concentrates local supervision where outcome feedback is weakest, with the caveats that it depends on the verifier and the teacher. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.