Yonghoon Dong, Minsung Yoon, Jaehyuk Kim, Jungwoo Park, Changyeon Kim, Jinwoo Shin
9 min
Abstract
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.
Alex: They examined the velocity Jacobian, a matrix describing how changes in an intermediate action affect the model’s velocity output. Averaged across a batch, that matrix was dominated by its diagonal. Cross-coordinate effects were relatively weak in that average.
Sam: A diagonal matrix can still treat different coordinates differently. That does not automatically give you one shared scaling factor.
Alex: That distinction matters. The authors go further and approximate the matrix with one shared scalar along its diagonal. Under that assumption, they prove that the backward signal always points along the same final-action gradient, with a changing magnitude.
Sam: What does the actual algorithm retain from that result, without requiring another expensive matrix calculation?
Alex: It drops an additional scaling factor that would require estimating the Jacobian again. What remains is the final-action gradient multiplied by flow time—how far action generation has progressed. Early steps get a weaker signal; later steps get a stronger one.
Sam: So the shortcut keeps one improvement direction and changes its strength. Do they check whether that resembles the original backward signal?
Alex: On the examined domains, both signals shrink similarly toward the beginning of generation and remain positively aligned. They are not identical. Their directions differ more on manipulation tasks, which are also where the shortcut alone performs poorly.
Sam: That is the obvious objection: cheaper training is not useful if it breaks the harder tasks. What fixes that failure?
Alex: The second ingredient controls the critic at the action the current policy actually generates. The shortcut uses that action’s gradient directly throughout the update. The authors hypothesize that critic errors there explain its failures.
Sam: Then the location of the regularization matters, not just whether the critic has some penalty attached.
Alex: Their value penalty discourages the critic from assigning a higher value to the generated action relative to the dataset action. It reuses the action already sampled for the policy update. There is no additional action-sampling cost.
Sam: Does keeping actions close to the dataset accomplish the same thing? That would also seem to limit unreliable extrapolation.
Alex: Not in their comparisons. An action-distance penalty constrains where the action lies, rather than the value assigned to it. Another alternative learns only from dataset actions. Neither directly controls the critic at the generated action whose gradient drives this update.
Sam: Could they just shrink that gradient instead? It is the signal the shortcut relies on.
Alex: They test that too. Penalizing the gradient’s magnitude helps, especially on the hardest cube manipulation domain. But the value penalty performs best on both cube domains where the unregularized shortcut struggles.
Sam: This makes the contribution a package: a cheaper update plus a critic trained where that update needs it.
Alex: And the value penalty helps the exact-adjoint method too, though less in the tested comparisons. So these experiments support special sensitivity to critic learning under the shortcut. They do not show that critic regularization matters only for this method.
Sam: Let’s put the performance claim in context. How broad was the main evaluation?
Alex: They evaluated fifty tasks in OGBench, covering navigation, manipulation, and planning. Each method started from the same behavior-cloned flow policy. Reported success rates average over eight random seeds.
Sam: What is the headline comparison, and does it hide tasks where the new method loses?
Alex: Average offline success reached seventy-five percent, against sixty-eight percent for the strongest baseline, trust-region adjoint matching. Gains concentrated on the hardest domains. On cube-quadruple, the advantage over the best baseline was thirty-five percentage points.
Sam: Concentrated gains are different from a uniform improvement. Where should a reader be cautious?
Alex: It trails the strongest baseline on the larger puzzle domain and on cube-double. The value penalty actually lowers cube-double’s offline performance, so the main experiments switch it off there. The method also retains a constraint on deviation from the pretrained policy for training stability.
Sam: And beyond simulation, what did the robot experiment actually test?
Alex: They fine-tuned a pretrained vision-language-action policy, which uses images and instructions to produce actions. A bimanual robot flipped a plastic bag, placed a straw in a cup, and put fruit in a pot before closing it. They trained each task separately, updating the policy’s action head.
Sam: How much evidence supports the improvement over supervised fine-tuning?
Alex: Each checkpoint was evaluated over thirty held-out episodes per task. For bag flipping, successes rose from fifteen with supervised fine-tuning to twenty-five after offline and online reinforcement learning. The method improved on the other robot tasks too.
Sam: Does that establish an advantage over residual corrections generally?
Alex: Only in this setting. The residual baseline struggled, but its released recipe was modified for the shared protocol. Initial configurations were fixed, and success was operator-judged. I would treat this as evidence of feasibility, not broad robot generalization.
Sam: What remains unresolved, and who should spend time with the full paper?
Alex: The authors cannot yet explain why the batch-averaged Jacobian is diagonal-dominated, or establish when that structure will hold. Flow-policy researchers should start with the scalar-adjoint derivation and the value-learning comparisons. Robot researchers should then read the evaluation protocol, especially the baseline modifications.
Sam: And what should everyone else remember?
Alex: A cheaper policy update can be enough—but its evaluator needs discipline at the actions the policy actually chooses.
Sam: Thanks for listening.