ResearchPod Summary
This paper investigates whether deep denoising models, which are known to produce human-like visual illusions in their final outputs, actually encode these illusions within their internal layers. While previous research has focused on output-level metrics like pixel intensity or classification decisions, this study uses fine-grained causal tracing to map where and how illusion-sensitive information is processed. The authors examine nine models across three architectures—pixel-space DDPMs, latent diffusion U-Nets, and diffusion transformers—to determine if the denoising objective drives the development of these internal representations.
The researchers identify that denoising models consistently develop illusion-sensitive representations at specific internal layers, particularly within the bottleneck of U-Net architectures. These internal activations are not merely statistical artifacts; they track validated human psychophysical models (such as FLODOG) and scale monotonically with the parametric strength of illusions like the Ebbinghaus and Ponzo effects. Through targeted channel ablation, the authors provide causal evidence that these specific channels significantly shape internal signals. However, they also discover a phenomenon they term 'perceptual phantoms': despite being causally active during internal processing, these representations do not propagate to the final image output. Injecting these internal illusion signals into a generation pipeline results in no measurable pixel shift, indicating a stark dissociation between internal perceptual processing and final output generation.
This work bridges the gap between 'black-box' output analysis and mechanistic interpretability in generative vision models. By demonstrating that denoising models develop human-like internal perceptual structures, the authors suggest that these illusions are a fundamental consequence of learning natural scene statistics rather than a byproduct of specific architectures. The identification of 'perceptual phantoms' provides a critical warning for researchers: internal representations that appear highly relevant to human perception may be completely absent from behavioral or output-based evaluations, necessitating more sophisticated internal-to-internal causal methods for future interpretability research.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.