ResearchPod Summary
Diffusion unlearning aims to remove harmful or copyrighted content from text-to-image models. Existing methods either use no anchor (anchor-free) or manually chosen anchors to guide the unlearning process. However, these methods often suffer from unstable updates that push the model away from the valid data manifold, leading to degraded image quality and incomplete concept removal. This paper investigates how to mathematically formalize this instability and proposes a method to generate more effective, manifold-proximal anchors.
The authors analyze diffusion unlearning through the lens of the manifold hypothesis, proving that without a manifold-proximal anchor, the update direction inevitably drifts into the normal space of the data manifold, causing unrobust unlearning. To address this, they introduce AutoAnchor, a two-stage framework. First, it automatically generates and filters a set of candidate concepts to serve as a semantic alternative to the target. Second, it optimizes these candidates using a novel cross-attention consistency loss, which acts as a computationally efficient surrogate for geometric manifold proximity. This allows the model to stay on the data manifold during the unlearning update.
Theoretical analysis confirms that existing unlearning methods are prone to significant normal-space drift. Empirical results demonstrate that AutoAnchor consistently outperforms state-of-the-art baselines, achieving up to a 31.04% improvement in CLIP scores for target concept removal and a 4.18% improvement in utility for non-target concepts. Furthermore, the framework is modular and can be integrated into existing unlearning pipelines, yielding an average performance boost of over 6% in both removal and utility metrics.
This work provides a rigorous mathematical foundation for why current diffusion unlearning techniques often fail to balance concept erasure with model performance. By automating the creation of stable anchors, AutoAnchor removes the need for manual, trial-and-error anchor selection, making the unlearning process more scalable, robust, and reliable for real-world applications where safety and utility must coexist.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.