ResearchPod Summary
Heterogeneous Knowledge Distillation (HKD) involves transferring knowledge from a high-capacity teacher (e.g., a Transformer) to a lightweight student (e.g., a CNN). This process is notoriously unstable due to two primary factors: massive discrepancies in feature magnitudes (structural gap) and conflicting gradient directions between the primary classification task and the distillation objective (optimization gap). The authors propose SPOFA (Spatial Projector and Momentum-based Adaptive One-For-All), a framework that addresses these issues through a dual-stabilization mechanism.
To address the structural gap, the authors introduce a LayerNorm-based decoupling projector. By applying Layer Normalization to the student's intermediate features before distillation, the framework explicitly separates feature magnitude from semantic direction. This prevents the student from wasting capacity on matching the teacher's absolute feature scales, allowing it to focus on aligning semantic representations.
To address the optimization gap, the authors introduce a Momentum-driven Exponential Moving Average (MEMA) dynamic scaler. Unlike existing methods that rely on volatile, instantaneous metrics to weight distillation losses, MEMA maintains a historical baseline of the optimization trajectory. It actively monitors the cosine similarity between primary and distillation gradients; when it detects that distillation signals are causing harmful conflicts, it adaptively penalizes those signals to ensure a smoother, more stable convergence.
Existing HKD methods often rely on computationally expensive attention modules or memoryless heuristics that fail to resolve the root causes of training instability. SPOFA achieves state-of-the-art accuracy on mainstream benchmarks like ImageNet and CIFAR while maintaining a very lightweight parameter footprint. By providing a principled way to decouple feature geometry and regulate gradient conflicts, the framework offers a more efficient and robust path for deploying high-performance models on resource-constrained devices.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.