ResearchPod Summary
Deep learning models are often black boxes, and existing methods for interpreting them face a difficult trade-off: post-hoc methods (like NMF) can be faithful to the model's actual decision process but produce unnamed, abstract factors, while by-design methods (like Concept Bottleneck Models) provide human-readable names but require retraining or altering the model. The authors ask: can we extract named, faithful, and spatially precise concepts from a frozen, pre-trained model without any modification?
The authors propose Language-Anchored Decomposition (LAD). Instead of learning both the concept basis and the coefficients (as in standard NMF), LAD fixes the coefficients using language-grounded similarity maps derived from CLIP. For each class, a large language model generates a vocabulary of relevant concepts (e.g., "pointy ears" for a cat). These concepts are localized across the image using CLIP-based similarity, creating a fixed coefficient matrix. LAD then learns only the concept basis that best reconstructs the frozen encoder's activations. Because the coefficients are fixed to language-anchored maps, the resulting basis vectors are inherently tied to human-interpretable names.
LAD successfully produces spatially precise, named concepts that are decision-relevant, as confirmed by concept insertion and deletion tests. Unlike prior methods that produce unnamed factors, LAD provides consistent, human-interpretable labels across different images. The authors demonstrate that the language anchor is not merely cosmetic; removing it preserves classification accuracy but significantly degrades attribution faithfulness, confirming that the language supervision steers the model toward discovering features that are genuinely meaningful to the classifier. Furthermore, LAD is computationally efficient, requiring only a single non-negative least-squares solve to fit the basis.
LAD bridges the gap between faithfulness and interpretability. By allowing researchers to explain the decisions of existing, high-stakes models without retraining, it provides a practical tool for diagnosing model failures and verifying that a model is relying on appropriate visual cues rather than spurious correlations. Its ability to provide stable, named concepts across different images makes it particularly useful for auditing models in fields like medical imaging or autonomous systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.