Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari
6 min
Abstract
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.
Alex: Take the Rubik's Cube case. They have to read cube state out of generated pixels, with no known camera pose. Doesn't the parsing introduce a lot of error?
Sam: That's the most significant constraint. Recovering 3D dynamics from 2D pixels is noisy. They fit a 3D cube model to the silhouette before checking facelet colours, and they manually verified samples to confirm the diagnostic overlays were accurate. But they acknowledge tracking failures are inevitable, so the metric is only as reliable as the parser.
Alex: So the symbolic evaluation is the load-bearing evidence. If the parser misreads a state, the validity count for that rollout is compromised. Is that why they use distributional metrics like CD-FVD elsewhere?
Sam: Yes. For datasets like MOVi-C or real-world video, you can't write a closed-form symbolic rule, so they move to distributional metrics. Those capture temporal distortions a parser might miss, but they don't give you the interpretability of rule-violation counts. You end up with two tiers: symbolic precision in controlled settings, distributional alignment in unconstrained ones.
Alex: Then the block-size ablation is supporting evidence rather than a headline. Why does smaller block size help on discrete datasets like Chess but barely matter for the double pendulum?
Sam: The explanation is about the nature of the dependencies. In a discrete state space, one invalid move in Chess cascades into total failure. Smaller blocks force more frequent causal commitments, which keeps the rollout on the rails. Continuous systems follow smooth, global laws, so larger blocks can still approximate the trajectory. The velocity field is less prone to sudden jumps. Note that this is an interpretation of the pattern, not something the ablation isolates directly.
Alex: Is the model learning different internal representations for each, or is it the same mechanism responding to different data?
Sam: The paper's framing is the latter. It isn't switching architectures. Block-causal attention is task-agnostic, and training shapes the weights toward the dependencies that matter in each regime.
Alex: And distribution shift? Chess is a clean case where standard models fail.
Sam: The authors show strong performance on symbolic tasks like Chess where standard models fail. The model still has to learn the transition between regimes, though, and I wouldn't extrapolate beyond the tasks they evaluated.
Alex: Even so, it nudges the scaling debate. If a change to how generation is scheduled reduces physical errors, some current failures may reflect how we force models to generate, not a lack of scale.
Sam: That's a fair reading, as long as it stays a suggestion. The evidence supports the claim on these benchmarks, with the parser caveat attached to the strongest ones.
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.