ResearchPod Summary
Post-hoc Concept Bottleneck Models (post-hoc CBMs) aim to make deep learning models interpretable by projecting latent features onto human-understandable concepts. However, researchers often rely on downstream classification accuracy to validate these models. This paper investigates whether such accuracy is a valid proxy for the semantic faithfulness of the learned concepts. The authors analyze two common training paradigms: using auxiliary concept datasets and using surrogate labels generated by vision-language models (VLMs). They formalize the conditions for faithfulness and propose new metrics to detect when concept projections fail to align with ground-truth semantics.
The study demonstrates that high downstream accuracy is insufficient to guarantee concept faithfulness. Using the manifold version of the Johnson-Lindenstrauss lemma, the authors show that even random, semantically meaningless concept projections can preserve enough information to allow a downstream classifier to achieve competitive performance.
Furthermore, the authors identify two critical sources of unfaithfulness:
This work challenges the common practice of using task accuracy as a primary evaluation metric for interpretability. By providing formal definitions and practical metrics for faithfulness, the authors offer researchers a way to verify whether their concept-based models are truly transparent or merely masking opaque decision-making processes. This is essential for building trust in AI systems used in high-stakes domains.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.