ResearchPod Summary
How can structural MRI, functional PET scans, and clinical assessment data be effectively combined to improve the accuracy and interpretability of Alzheimer's disease (AD) classification? The authors seek to overcome the limitations of unimodal models and the black-box nature of deep learning in clinical diagnostics.
The researchers developed a multimodal framework using ResNet50 backbones to extract features from 2D MRI and PET slices. These imaging features are aggregated at the patient level and concatenated with clinical variables (such as MMSE scores and APOE4 status) into a final fusion network. To address the black-box problem, the authors applied Gradient-weighted Class Activation Mapping (Grad-CAM) to visualize which brain regions influence the model's predictions, providing clinicians with a transparent diagnostic aid.
The study demonstrates that integrating clinical data with neuroimaging provides the most significant boost to diagnostic performance. In an ablation study, the model's accuracy improved from 39.61% (MRI only) and 51.72% (PET only) to 79.17% when all three modalities were combined. The authors also show that the Grad-CAM heatmaps for PET scans are class-sensitive, meaning they change based on the diagnostic hypothesis, which supports the model's potential for clinical decision support.
This work highlights that clinical and cognitive assessment data contain diagnostic signals that are largely complementary to neuroimaging. By providing a framework that is both multimodal and explainable, the authors offer a path toward more transparent AI tools that clinicians can trust, though they emphasize that larger, multi-center validation is required to confirm these results for real-world deployment.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.