Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Sam: Updating diffusion models sparsely, and early in the denoising process, appears to raise reward without the diversity collapse that late-timestep fine-tuning tends to produce. That comes from a paper introducing a method called iADD.
Alex: That cuts against my assumption. I thought reward signals were more stable when backpropagated through the final denoising steps. Why would early steps be the better target?
Sam: Late steps are stable, but they have little leverage, because by then the structure is already committed. Early timesteps have broader generative reach, so updates there steer the sample before the details are fixed. The authors pair this with Feynman-Kac path guidance. Several particles run through denoising, intermediate potentials score them, and low-potential paths are pruned early.
Alex: So rather than pushing the final output toward a reward, you're curating the whole trajectory. Bad branches get dropped before they commit.
Sam: Roughly, yes. The curriculum is also sparse: at each stage only a subset of timesteps is updated and the rest stay frozen. The idea is that this avoids over-constraining the model, and over-constraining is what drives mode collapse when you force reward at the expense of the original distribution.
Alex: Let's anchor that in a task. In the 3D indoor scene synthesis experiments, Table 3 has the DDPO baseline at about seven point seven percent collisions, against about six point one for iADD. That's a clear improvement, but a modest one. Why would a sparse curriculum specifically reduce object overlap?
Sam: The paper's explanation is about where the layout gets decided. Early steps set the global arrangement, such as where the bed sits relative to the TV stand. If you update only the late steps, the model can learn a shortcut: put objects in fixed, high-probability locations that satisfy the reward, and ignore the spatial constraints. You end up trying to repair a bad floor plan by shuffling furniture at the last second, and that's where the collisions come from.
Alex: So early steps act as a global planner and late steps do local refinement. But that's an interpretation. What's the more direct evidence?
Sam: The closest thing is the Jacobian log-determinant analysis in Figure 6. Dense updates give a markedly more negative log-determinant, which signals volume contraction, in other words mode collapse. Keeping most timesteps frozen keeps that volume intact. It lines up with the Inception Score staying higher as reward rises. But it's a diagnostic consistent with the account, not a direct test of the global-planner reading, and I'd keep that distinction in mind.
Alex: And the cost? Particle-based sampling sounds expensive.
Sam: It is. To reach a given CLIP reward target, iADD needs roughly seven point four GPU hours, against about two point five for DDPO. The overhead comes from scoring multiple particles during rollout. The incremental schedule trims some of it by reducing backward passes early on, but particle branching is a structural choice. The authors defend it as the price of avoiding the reward-diversity trade-off.
Alex: One more design choice stood out. Most RL-based diffusion fine-tuning leans on explicit KL regularization, and this doesn't.
Sam: Right. Methods like DPOK use a KL term to hold diversity in place. The authors' position is that sparse, incremental updates preserve the prior's probability volume on their own, so the constraint isn't needed. The training schedule does the work a regularizer usually patches in afterwards. Whether that holds beyond the settings they tested is the open question.
Alex: So the trade is roughly threefold compute for a structural fix rather than a loss-function patch, with a modest collision gain on this one task. Whether that's worth it probably depends on how much you care about preserving diversity.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.