ResearchPod Summary
Building foundation models for medical imaging is hindered by privacy regulations that prevent centralized data pooling. Existing federated approaches often struggle with Imaging Modality Heterogeneity, where institutions operate with different imaging hardware (e.g., MRI vs. CT) and non-IID label distributions. The authors investigate how to build a unified, multi-task federated foundation model that can handle both overlapping and disjoint modality configurations without requiring raw data exchange.
The authors propose FM2, a framework that trains a core visual backbone from scratch to maintain medical domain fidelity. The architecture features two key innovations:
FM2 demonstrates consistent superiority over state-of-the-art federated baselines across classification, caption-supervised learning, and medical Visual Question Answering (VQA). The authors provide formal proofs showing that the framework achieves a convergence rate of O(1/√T) and explicit generalization bounds. Experimental results on the newly constructed MIMH benchmark confirm that the dual-MoE structure effectively disentangles class-level personalization from modality-level consensus, enabling strong out-of-modality generalization.
This work provides a scalable solution for clinical AI, allowing hospitals to collaborate on foundation models without compromising patient privacy or requiring uniform imaging hardware. By treating language as a semantic bridge, FM2 enables institutions with vastly different data types to contribute to a unified, high-performance medical AI ecosystem.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.