Ibrahim, Irwansyah and Alshar'e, Marwan and Sanjaya, Imam and Tiwari, Gagan (2026) Cross-Modal Representation Learning for Integrating Heterogeneous Data in AI Systems. Journal of Data Science, 2026 (13). pp. 218-235. ISSN 2805-5160
|
Text
jods2026_13.pdf - Published Version Available under License Creative Commons Attribution. Download (424kB) |
|
|
Text
931 - Published Version Available under License Creative Commons Attribution. Download (40kB) |
Abstract
The increasing availability of heterogeneous data sources, including text, images, and structured records, has intensified the need for robust multimodal artificial intelligence systems. However, existing multimodal learning approaches often rely on simplistic fusion strategies and struggle to capture deep semantic relationships across modalities, leading to limited robustness, poor representation consistency, and reduced performance under incomplete data conditions. To address this gap, this study proposes a cross-modal representation learning framework that aligns heterogeneous modalities within a shared latent representation space. The proposed framework integrates modality-specific encoders, contrastive alignment learning, distribution alignment constraints, and attention-based fusion to enable semantically coherent and adaptive multimodal interaction. Experiments were conducted on multiple multimodal benchmark datasets using repeated evaluation settings and standard performance metrics, including accuracy, precision, recall, and F1-score. The results demonstrate that the proposed framework consistently outperforms unimodal and conventional fusion methods, achieving the best classification accuracy of 91.6% and an F1-score of 90.9%. Furthermore, the framework exhibits strong robustness under missing modality conditions, with significantly lower performance degradation compared to baseline approaches. Latent space analysis and ablation studies further confirm the effectiveness of cross-modal alignment in improving representation consistency and generalization capability. The primary goal of this research is to develop a scalable, interpretable, and resilient framework for integrating heterogeneous data in complex Al environments. The findings contribute to advancing multimodal representation learning for next-generation intelligent systems.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Cross-Modal Learning, Multimodal AI, Representation Learning, Heterogeneous Data Integration, Attention-Based Fusion |
| Subjects: | Q Science > QA Mathematics Q Science > QA Mathematics > QA75 Electronic computers. Computer science Q Science > QA Mathematics > QA76 Computer software |
| Depositing User: | Unnamed user with email masilah.mansor@newinti.edu.my |
| Date Deposited: | 11 Sep 2026 07:07 |
| Last Modified: | 11 Sep 2026 07:07 |
| URI: | http://eprints.intimal.edu.my/id/eprint/2368 |
Actions (login required)
![]() |
View Item |
