FedCAD: Cross-modal semantic alignment and distillation for cross-domain heterogeneous federated learning.
Where this comes from
- Record sourced from PubMed, PMID 42372644.
- Also identified by DOI 10.1016/j.neunet.2026.109298.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recent advances in the IoT and edge intelligence have made the deployment of CLIP-style image-text models in edge-cloud architectures increasingly common for collaborative sensing. However, resource heterogeneity at the edge limits the feasibility of using a unified backbone across devices with different computational budgets. In addition, non-IID data and domain shifts can disrupt image-text alignment and cause semantic drift during federated aggregation. To address these challenges, we propose FedCAD, a federated learning framework for effective knowledge collaboration under cross-domain distribution shifts and heterogeneous edge environments through cross-modal semantic alignment and multi-level distillation. FedCAD consists of three main components: (i) a Latency-Distribution Co-aware Clustering (LDCC) strategy with heterogeneous model orchestration to alleviate resource disparities and straggler effects; (ii) a decouple-then-align dual-stage feature alignment mechanism that suppresses domain-specific noise and preserves the geometric consistency of the joint image-text embedding space, thereby enhancing cross-domain generalization; and (iii) a multi-stage collaborative distillation protocol spanning intra-cluster, inter-cluster, and cloud levels, which promotes cross-cluster semantic complementarity and cross-architecture knowledge fusion to mitigate knowledge fragmentation. Experimental results show that FedCAD consistently improves classification accuracy and target-domain generalization across multiple cross-domain classification benchmarks. In addition, results on Flickr30K and MSCOCO further demonstrate its effectiveness in image-text matching and cross-modal retrieval. Under the frozen-backbone and lightweight-adapter setting, only adapter parameters are transmitted during communication rounds, which to some extent supports its deployment applicability in bandwidth-constrained networks.