Ramil Khafizov, Ilya Statsenko, Ruslan Rakhimov, Artem Komarichev, Peter Wonka, Evgeny Burnaev
9 min
Abstract
Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/
Alex: So it keeps sequential dependence between coarse and fine stages, but avoids a long sequence over every image position.
Sam: At each stage, it predicts visual tokens—compact pieces of an image representation—in parallel. It also predicts that stage across all requested target views in parallel. Later stages depend on earlier stages, so this is not a single-pass generator.
Alex: Parallel target views could still drift apart. What actually connects them while they’re being generated?
Sam: Attention lets the model exchange information between image features. At a given scale, target views can communicate freely with each other and consult earlier scales. They cannot consult future scales, which preserves the coarse-to-fine generation order.
Alex: That explains communication, but not geometry. How does the model know which views belong to which cameras?
Sam: The authors add Multi-scale Projective Pose Encoding. It puts camera transformations into attention while retaining each token’s local image position. They adapt that encoding to the changing grids used at successive scales.
Alex: Is camera information only attached to the input, or does it shape the exchanges throughout generation?
Sam: It shapes both major exchanges throughout generation. Communication among target views uses their target cameras. When target features consult source features, the two sides use their respective cameras, anchoring generation to the observed images.
Alex: Give me a concrete detail that distinguishes this from simply handing the model a camera label.
Sam: Even the coarsest target stage can consult the full-resolution source feature map. The source evidence isn’t reduced to the target’s current coarse grid. So detailed appearance information remains available from the beginning, through geometry-aware attention.
Alex: There’s also an overall object identity to preserve, not just local detail. Does the model get a separate signal for that?
Sam: Yes, through a complementary global path. It pools source features into a summary that starts generation and modulates the transformer blocks. The dense path keeps access to unpooled source features; the design aims to preserve both overall appearance and spatial detail.
Alex: What makes this genuinely new rather than another autoregressive wrapper around diffusion?
Sam: Its image generation is diffusion-free. The paper contrasts it with CausNVS, which uses an autoregressive sequence but retains diffusion steps for individual views. NAMVIS instead combines discrete coarse-to-fine prediction with camera-aware communication between views.
Alex: Let’s move to evidence. What did they train on, and how broadly did they test it?
Sam: They curated approximately two hundred thousand objects from Objaverse-XL. Filtering removed rendering failures, poor textures, problematic geometry, and other low-quality assets. They tested on Objaverse, Google Scanned Objects, and OmniObject3D, with varying source and target view counts.
Alex: If the outside datasets also improve, that supports transfer beyond the training collection. What does “improve” mean here?
Sam: They measure pixel fidelity, structural similarity, and perceptual similarity against ground-truth renders. NAMVIS beats the evaluated diffusion baselines on all those measures across each dataset. Performance still drops outside Objaverse, but the authors report more graceful degradation than the compared methods.
Alex: What’s the speed difference under the actual test conditions?
Sam: In the single-source, single-target setting, NAMVIS takes point six seconds per target view. The fastest evaluated diffusion baseline, Wonder Three D, takes two seconds. That comparison uses images two hundred fifty-six pixels on each side, not high-resolution output.
Alex: Were the baselines all operating under identical camera conditions? Those restrictions could affect how we interpret the ranking.
Sam: They used official pretrained weights, but followed each method’s supported camera protocol. Wonder Three D uses a different elevation from the other methods. For single-view baselines, they also select the available source giving the best pixel-fidelity result, so this isn’t one completely uniform protocol.
Alex: And better individual images don’t establish a coherent object. What evidence addresses that gap?
Sam: They use COLMAP, software that reconstructs cameras and sparse geometry from overlapping images. Consistent views should support feature matching and triangulation more readily. NAMVIS produces more reconstructed points than the compared multi-view diffusion methods, although its lead over EscherNet is modest.
Alex: Can they also check whether the generated cameras are successfully registered, rather than only counting reconstructed points?
Sam: Across thirty scenes, NAMVIS registers about sixty-eight percent of views, versus about sixty percent for EscherNet. Ground-truth images reach seventy percent. That supports consistency, but reconstructability remains a proxy rather than proof of correct geometry.
Alex: Do the component tests support the proposed mechanism, and what remains untested?
Sam: Smaller-data ablations favor the camera encoding and show worse results without dense source attention. But the current checkpoint requires more training for higher resolution. The larger uncertainty is transfer from object-centric data to complex scenes; the authors identify that as future work, alongside robustness to pose noise.
Alex: Given that boundary, who should read the full document, and where should they start?
Sam: Researchers building generative novel-view systems should start with the Method section, especially camera encoding and image conditioning. Then read Baselines to understand the comparison protocols, and Limitations and future work before planning deployment. For everyone else: coarse-to-fine generation can replace diffusion without giving up communication between camera views—but the evidence here is for objects at modest resolution.
Alex: Thanks for listening.