Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
4 min
Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Sam: It is. To reach a given CLIP reward target, iADD needs roughly seven point four GPU hours, against about two point five for DDPO. The overhead comes from scoring multiple particles during rollout. The incremental schedule trims some of it by reducing backward passes early on, but particle branching is a structural choice. The authors defend it as the price of avoiding the reward-diversity trade-off.
Alex: One more design choice stood out. Most RL-based diffusion fine-tuning leans on explicit KL regularization, and this doesn't.
Sam: Right. Methods like DPOK use a KL term to hold diversity in place. The authors' position is that sparse, incremental updates preserve the prior's probability volume on their own, so the constraint isn't needed. The training schedule does the work a regularizer usually patches in afterwards. Whether that holds beyond the settings they tested is the open question.
Alex: So the trade is roughly threefold compute for a structural fix rather than a loss-function patch, with a modest collision gain on this one task. Whether that's worth it probably depends on how much you care about preserving diversity.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.