Kaiyue Wen, Luke Bailey, Arvind Mahankali, Tengyu Ma
5 min
Standard reinforcement learning for LLMs, such as Group Relative Policy Optimization (GRPO), typically relies on terminal rewards to evaluate entire reasoning trajectories. This coarse credit assignment penalizes potentially correct reasoning if a minor error occurs at the end. The authors propose Actor-Critic with Action Chunking (AC2), which treats rollouts as sequences of action chunks. By using a learned critic to score the state at the end of each chunk, the policy can update without waiting for a terminal reward. To ensure the critic remains reliable, AC2 employs three design choices: local readiness (updating only when the critic is sufficiently accurate on a specific problem), using reference solutions to guide the critic, and evaluating chunks of 10,000 tokens to provide a meaningful signal.
AC2 significantly improves compute efficiency and performance compared to GRPO. On the IMO-ProofBench benchmark, AC2 exceeds GRPO's peak validation score of 18.5% while using 2.5 times fewer decoding FLOPs. This efficiency gain stems from two factors: AC2 requires 25% fewer training steps to reach peak performance, and it generates fewer tokens per update because it does not need to complete every trajectory. The authors demonstrate that the critic becomes trustworthy early in training and that the combination of local readiness, reference solutions, and action chunking is essential for these gains.
This work challenges the prevailing view that learned critics are too inaccurate for LLM reinforcement learning. By demonstrating that LLMs can effectively trust a learned critic to perform credit assignment on partial trajectories, the authors open a new design space for RL algorithms. This shift allows for more flexible training strategies, such as learning from off-policy initial states and performing more efficient, fine-grained updates, which are critical for scaling reasoning capabilities in models.
Sam: That also suggests a way to handle distribution shift. Say the model meets problems much harder than the training set.
Alex: The gate is dynamic, so it should handle that. If new problems push critic error above threshold, they stop counting as ready and revert to full rollouts until the critic learns to value those states. Compute scales with how far the critic can currently be trusted. I would note that this is the mechanism as described; it isn't the same as a dedicated shift experiment.
Sam: Which ablation matters most for the headline comparison?
Alex: The ones around the gate. Removing the local readiness check produces a stark degradation. The model is effectively trusting a critic that hasn't earned it.
Sam: I noticed that the variant without group-and-audit, but with local readiness, still beats the one with readiness disabled entirely. That suggests the gate is doing most of the work, not the auditing.
Alex: That is how I read it. Even without the full auditing loop, the policy avoids updating on states where the critic is demonstrably uncertain, so it doesn't drift into poor regions. The auditing then adds a further safeguard on top of that.
Sam: It's an interesting inversion of PPO. There, the clipped surrogate objective keeps updates stable. Here, stability comes from the trustworthiness of the value function, and the critic has to prove its accuracy before the policy may use it.
Alex: Yes. The critic stops being a passive baseline and becomes an active guide, but only on problems where it has demonstrated it can be one.
Sam: What would a careful referee push on?
Alex: The main limitation is data regime. Calibrating the gate requires periodic ground-truth values from full rollouts, which means epoching over the training data. That makes it hard to use in a single-epoch, streaming setting, so it is batch-oriented by design.
Sam: So it addresses the rollout cost for verifiable reasoning tasks, but it may not suit every RL pipeline. Still, much of the compute tax may come from not trusting our value functions, and a reliable gate changes that.
Alex: That's how I'd frame it. The evidence suggests full-trajectory rollouts are not always necessary if you can tell when the critic is reliable, and that opens a design space that has not been explored much.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.