ResearchPod Summary
How can generative models be audited for training data membership without relying on computationally expensive shadow models or fragile one-shot queries? The authors investigate whether the dynamics of recursive self-generation can serve as a signal amplifier to expose whether specific samples were used during training.
Inspired by the phenomenon of Model Autophagy Disorder (MAD), the authors propose MADreMIA, an inference-time framework that creates iterative trajectories for queried samples. By feeding the output of a model back into itself as the input for the next step, the framework generates a sequence of outputs. The authors hypothesize that memorized training samples act as stable attractors in the model's latent space, leading to higher structural coherence and slower semantic degradation compared to non-member samples, which drift toward the model's average bias or dissolve into noise.
As large generative models are increasingly scrutinized for privacy and copyright compliance, traditional one-shot inference attacks often fail due to their sensitivity to distributional shifts. MADreMIA provides a scalable, model-agnostic tool for privacy auditing that leverages the inherent structural "echoes" of training data, offering a more reliable way to verify whether specific content was used to shape a model's parameters.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.