Home > Published Issues > 2026 > Volume 17, No. 9, 2026 >
JAIT 2026 Vol.17(9): 1667-1675
doi: 10.12720/jait.17.9.1667-1675

Selective Cross-modal Feature Enhancement through Collaborative Learning

Mohammed Belghadi 1,*, Younès Bennani 1, and Valérie Paradis 2
1. LIPN-CNRS, UMR 7030-LaMSN, Université Sorbonne Paris Nord, Paris, France
2. INSERM, UMR 1149-CNRS EMR 8252, Université Paris Cité, Paris, France
Email: mohammed.belghadi@sorbonne-paris-nord.fr (M.B.); younes.bennani@sorbonne-paris-nord.fr (Y.B.); valerie.paradis@aphp.fr (V.P.)
*Corresponding author

Manuscript received January 9, 2026; revised February 12, 2026; accepted May 25, 2026; published September 4, 2026.

Abstract—Multimodal survival prediction is based on integrating heterogeneous data sources—including radiology, pathology, genomics, and clinical attributes—that are often incomplete in real clinical workflows. Most existing approaches depend on late fusion, which limits cross-modal interaction and degrades under missing modalities. We introduce Selective Cross-Modal Learning (SCML), a lightweight pre-fusion refinement framework that enriches unimodal embeddings by selectively incorporating complementary cues from concurrently available modalities. SCML operates in three stages: (1) projection-space alignment to enable meaningful cross-modal similarity measurement with-out corrupting task-discriminative representations; (2) selective cross-modal attention guided by pairwise synergy analysis and global modality-importance weighting, ensuring that each modality is enhanced only by empirically supportive counterparts; and (3) a reconstruction constraint that preserves modality-specific structure and improves robustness to incomplete data. We evaluated SCML in a cohort of 962 patients with glioma using 15-fold Monte Carlo cross-validation, achieving a median C-index of 0.8147—surpassing strong baselines including Mean-Vector fusion, Pathomic Fusion, and Multimodal learning with Missing Data (MMD). SCML maintains its advantage under missing-pathology and missing-genomics conditions. The framework is encoder-agnostic, modality-agnostic, and readily applicable to other multimodal prediction tasks that involve partially missing and heterogeneous inputs.
 
Keywords—multimodal learning, cross-modal attention, missing modalities, feature refinement, representation learning

Cite: Mohammed Belghadi, Younès Bennani, and Valérie Paradis, "Selective Cross-modal Feature Enhancement through Collaborative Learning," Journal of Advances in Information Technology, Vol. 17, No. 9, pp. 1667-1675, 2026. doi: 10.12720/jait.17.9.1667-1675

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions