ResearchPod Summary
Post-hoc explanation methods typically evaluate model decisions by measuring prediction changes under altered inputs using perturbation or counterfactual techniques. While these methods measure how much a model reacts through response magnitude, magnitude alone fails to explain what that reaction means. An identical absolute response can represent a stable effect supporting the final difference, an opposing reaction, or a large intermediate change that vanishes at the endpoint. To resolve this ambiguity, the paper introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility), a model-agnostic method that routes paired reveal responses into semantic components.
The DECAF framework operates on paired reveal trajectories where a factual-counterfactual input pair is progressively revealed from a common uninformative state. Using the final contrast between the fully observed pair as a semantic reference, the method applies a practical threshold and an orientation sign. Active responses aligned with the final contrast are routed to evidence (E), opposed responses are routed to contradiction (C), and responses on final-contrast-null pairs are routed to fragility (F). This routing preserves ordinary magnitude exactly such that absolute response equals the sum of the three components.
Theoretical analysis establishes that this semantic routing is unique under endpoint-relative axioms and that DECAF strictly refines ordinary response magnitude. Across controlled vision and tabular experiments spanning multiple model architectures, the three components map directly to independently measured behaviors: evidence tracks shortcut reliance and learned feature presence, fragility tracks off-path sensitivity on endpoint-null pairs, and contradiction successfully isolates effect reversal from mere attenuation.
In a 72-model ImageNet-9 audit comparing cases with matched response magnitudes, DECAF's largest component aligns with observed behavior in 96.4% of cases compared to 35.0% for magnitude alone. Furthermore, short forward-only DECAF trajectories outperform standard attribution baselines on FunnyBirds and ImageNet-1k, and achieve competitive performance with significantly lower wall time and memory overhead on a 1-billion-scale DINOv2 model.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.