ResearchPod Summary
How does the process of memorizing noisy labels in deep neural networks manifest in the geometry of training gradients? While previous research has focused on loss dynamics or representation learning, this paper investigates the spectral properties of the centered Fisher-gradient scatter matrix to understand how memorization reshapes the optimization landscape.
The authors analyze the effective rank of the centered Fisher-gradient scatter matrix—the covariance of per-example gradients at the final layer. They define 'Fisher Rank Inflation' as the transient expansion of this effective rank during the memorization phase. To identify which specific training examples drive this inflation, they derive a first-order leave-one-out (LOO) attribution formula. This allows them to quantify how much each individual example contributes to the rank expansion. They validate these findings across multiple architectures (SmallCNN, ResNet18, Vision Transformers) and datasets (CIFAR-10, CIFAR-100, and CIFAR-10N), testing both synthetic symmetric label noise and naturally occurring human annotation errors.
The study establishes that Fisher Rank Inflation is a consistent phenomenon across diverse architectures and noise types. During training, the effective rank of the gradient scatter expands as the network begins to memorize corrupted labels, peaking during the memorization phase, and then collapsing once the corrupted labels are fully fit. At the peak of this inflation, corrupted examples are significantly enriched among the samples that contribute most to the rank increase. The authors also show that the peak effective rank grows monotonically with the severity of label corruption. Notably, the onset of this rank inflation often precedes observable degradation in test set performance, suggesting it could serve as an early warning signal for memorization.
This work provides a novel, geometry-based perspective on the 'early learning' phenomenon in deep learning. By linking the spectral structure of gradients to the memorization of noisy labels, the authors offer a theoretical framework that explains why certain examples are more influential during specific training phases. This spectral signature provides a more granular view of training dynamics than loss-based metrics, potentially enabling better diagnostic tools for identifying noisy data or monitoring the transition from meaningful learning to overfitting.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.