ResearchPod Summary
Federated Learning (FL) allows multiple entities to train machine learning models collaboratively without sharing raw data. However, standard FL implementations often rely on a central server that can access model updates, creating privacy risks and a single point of failure. The authors aim to build a flexible, secure, and scalable framework that supports threshold homomorphic encryption (ThHE) to enable blind model aggregation while maintaining modularity for diverse machine learning and cryptographic applications.
MOSAIC-FL utilizes a micro-service architecture where each node is composed of three isolated, specialized containers: an Orchestrator, an ML Engine, and a Crypto Provider. This design enforces a strict separation of concerns, where components communicate via gRPC and Protocol Buffers to minimize overhead while ensuring language-agnostic extensibility. The framework implements a threshold CKKS homomorphic encryption scheme, which allows the central server to aggregate encrypted model updates without ever accessing the plaintext weights. Decryption is only possible if a threshold of t-out-of-N participants cooperate, mitigating the risk of a malicious or compromised central server.
The authors demonstrate the framework's effectiveness through two primary use cases: standard image recognition (EMNIST) and complex genomic classification (breast cancer subtyping using TCGA data). The architecture successfully handles the computational and communication overhead associated with homomorphic encryption by using efficient gRPC protocols and renewing collective key material at every round to prevent key-recovery attacks. The results indicate that the framework is capable of performing secure, privacy-preserving training across different model scales and threshold configurations, providing a robust alternative to monolithic FL frameworks that lack native support for threshold cryptography.
By decoupling the cryptographic, communication, and machine learning logic, MOSAIC-FL addresses the rigidity of existing FL frameworks. Its ability to perform secure aggregation in sensitive domains like genomics—where data privacy is paramount—without requiring a fully trusted central authority makes it a significant contribution to privacy-preserving machine learning infrastructure.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.