Cross-Modal Representation Learning for Integrating Heterogeneous Data in AI Systems

Authors

  • Irwansyah Ibrahim Universitas Bina Darma, Palembang, Indonesia
  • Marwan Alshar'e Sohar University, Sohar, Oman
  • Imam Sanjaya Nusa Putra University, Sukabumi, Indonesia
  • Gagan Tiwari Noida International University, Uttar Pradesh, India

DOI:

https://doi.org/10.61453/jods.v20260213

Keywords:

Cross-Modal Learning, Multimodal AI, Representation Learning, Heterogeneous Data Integration, Attention-Based Fusion

Abstract

The increasing availability of heterogeneous data sources, including text, images, and structured records, has intensified the need for robust multimodal artificial intelligence systems. However, existing multimodal learning approaches often rely on simplistic fusion strategies and struggle to capture deep semantic relationships across modalities, leading to limited robustness, poor representation consistency, and reduced performance under incomplete data conditions. To address this gap, this study proposes a cross-modal representation learning framework that aligns heterogeneous modalities within a shared latent representation space. The proposed framework integrates modality-specific encoders, contrastive alignment learning, distribution alignment constraints, and attention-based fusion to enable semantically coherent and adaptive multimodal interaction. Experiments were conducted on multiple multimodal benchmark datasets using repeated evaluation settings and standard performance metrics, including accuracy, precision, recall, and F1-score. The results demonstrate that the proposed framework consistently outperforms unimodal and conventional fusion methods, achieving the best classification accuracy of 91.6% and an F1-score of 90.9%. Furthermore, the framework exhibits strong robustness under missing modality conditions, with significantly lower performance degradation compared to baseline approaches. Latent space analysis and ablation studies further confirm the effectiveness of cross-modal alignment in improving representation consistency and generalization capability. The primary goal of this research is to develop a scalable, interpretable, and resilient framework for integrating heterogeneous data in complex Al environments. The findings contribute to advancing multimodal representation learning for next-generation intelligent systems.

References

Bose, P., Rana, P., & Ghosh, P. (2023). Attention-Based Multimodal Deep Learning on Vision-Language Data: Models, Datasets, Tasks, Evaluation Metrics and Applications. IEEE Access, 11, 80624–80646. https://doi.org/10.1109/ACCESS.2023.3299877

Cabot, J. H., & Ross, E. G. (2023). Evaluating prediction model performance. Surgery, 174(3), 723-726. https://doi.org/10.1016/J.SURG.2023.05.023

Chen, J., Seng, K. P., Smith, J., & Ang, L. M. (2024). Situation Awareness in AI-Based Technologies and Multimodal Systems: Architectures, Challenges and Applications. IEEE Access, 12, 88779–88818. https://doi.org/10.1109/ACCESS.2024.3416370

Hu, Z., Gutiérrez-Basulto, V., Xiang, Z., Li, R., & Pan, J. Z. (2026). Leveraging intra-modal and inter-modal interaction for multi-modal entity alignment. Neurocomputing, 676, 133017. https://doi.org/10.1016/J.NEUCOM.2026.133017

Jabeen, S., Li, X., Amin, M. S., Bourahla, O., Li, S., & Jabbar, A. (2023). A Review on Methods and Applications in Multimodal Deep Learning. ACM Transactions on Multimedia 19(2s), Computing, Communications and Applications, https://doi.org/10.1145/3545572

Li, S., & Tang, H. (2026). Multimodal Alignment and Fusion: A Survey. International Journal of Computer Vision 2026 134:3, 134(3), 103-. https://doi.org/10.1007/S11263-025-02667-1

Li, Z., Yao, T., Wang, L., Li, Y., & Wang, G. (2024). Supervised Contrastive Discrete Hashing for cross-modal retrieval. Knowledge-Based Systems, 295, 111837. https://doi.org/10.1016/J.KNOSYS.2024.111837

Liang, X., Yang, E., Deng, C., & Yang, Y. (2024). CrossFormer: Cross-Modal Representation Learning via Heterogeneous Graph Transformer. ACM Transactions on Multimedia Computing, Applications, 20(12), Communications and https://doi.org/10.1145/3688801

Liao, L., Li, H., Shang, W., & Ma, L. (2022). An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks. and Methodology, 31(3). ACM Transactions on Software Engineering https://doi.org/10.1145/3506695

Lin, H., Zhang, C. Y., & Philip Chen, C. L. (2025). Contextual Distribution Alignment via Correlation Contrasting for Domain Generalization. IEEE Transactions on Circuits and 35(4), Systems for Video Technology, https://doi.org/10.1109/TCSVT.2024.3509902

Liu, H., Shi, Y., Li, A., & Wang, M. (2024). Multi-modal fusion network with intra- and inter-modality attention for prognosis prediction in breast cancer. Computers in Biology and Medicine, 168, 107796. https://doi.org/10.1016/J.COMPBIOMED.2023.107796

Manna, S., Chattopadhyay, S., Dey, R., Pal, U., & Bhattacharya, S. (2025). Dynamically Scaled Temperature in Self-Supervised Contrastive Learning. IEEE Transactions on Artificial Intelligence, 6(6), 1502–1512. https://doi.org/10.1109/TAI.2024.3524979

Manzoor, M. A., Albarri, S., Xian, Z., Arslan, M., Meng, Z., Nakov, P., Liang, S., Manzoor, M. A., Albarri, S., Nakov, P., & Liang, S. (2023). Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications. ACM Transactions on Multimedia Computing, Communications and Applications, 20(3), 74. https://doi.org/10.1145/3617833

Novac, O. C., Chirodea, M. C., Novac, C. M., Bizon, N., Oproescu, M., Stan, O. P., & Gordan, C. E. (2022). Analysis of the Application Efficiency of TensorFlow and PyTorch in Convolutional Neural Networks. Sensors 2022, Vol. 22(22). https://doi.org/10.3390/s22228872

Rajendran, S., Pan, W., Sabuncu, M. R., Chen, Y., Zhou, J., & Wang, F. (2024). Learning across diverse biomedical data modalities and cohorts: Challenges and opportunities for innovation. Patterns, 5(2), 100913. https://doi.org/10.1016/j.patter.2023.100913

Soenksen, L. R., Ma, Y., Zeng, C., Boussioux, L., Villalobos Carballo, K., Na, L., Wiberg, H. M., Li, M. L., Fuentes, I., & Bertsimas, D. (2022). Integrated multimodal artificial intelligence framework for healthcare applications. Npj Digital Medicine 2022 5:1, 5(1), 149-. https://doi.org/10.1038/s41746-022-00689-4

Tan, K., Huang, W., Liu, X., Hu, J., & Dong, S. (2022). A multi-modal fusion framework based on multi-task correlation learning for cancer prognosis prediction. Artificial Intelligence in Medicine, 126, 102260. https://doi.org/10.1016/J.ARTMED.2022.102260

Tian, J., Xiong, R., Shen, W., Lu, J., & Sun, F. (2022). Flexible battery state of health and state of charge estimation using partial charging data and deep learning. Energy Storage Materials, 51, 372–381. https://doi.org/10.1016/J.ENSM.2022.06.053

Wang, T., Li, F., Zhu, L., Li, J., Zhang, Z., & Shen, H. T. (2024). Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions. Proceedings of the IEEE, 112(11), 1716-1754. https://doi.org/10.1109/JPROC.2024.3525147

Wu, X., Shi, Y., Wang, M., & Li, A. (2023). CAMR: cross-aligned multimodal representation learning for cancer survival prediction. Bioinformatics, 39(1). https://doi.org/10.1093/BIOINFORMATICS/BTAD025

Zha, D., Bhat, Z. P., Lai, K. H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2025). Data-centric Artificial Intelligence: A Survey. ACM Computing Surveys, 57(5), 129. https://doi.org/10.1145/3711118

Zhang, Z., Yu, C., Zhang, H., & Gao, Z. (2024). Embedding Tasks Into the Latent Space: Cross-Space Consistency for Multi-Dimensional Analysis in Echocardiography. IEEE 2215-2228. Transactions on Medical Imaging, 43(6), https://doi.org/10.1109/TMI.2024.3362964

Zhao, M., Yuan, Y., Luo, L., & Li, X. (2025). A Review: Absolute Linear Encoder Measurement Technology. Sensors 2025, Vol. 25, 25(19). https://doi.org/10.3390/S25195997

Zhou, H. Y., Yu, Y., Wang, C., Zhang, S., Gao, Y., Pan, J., Shao, J., Lu, G., Zhang, K., & Li, W. (2023). A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics. Nature Biomedical Engineering 2023 7:6, 7(6), 743-755. https://doi.org/10.1038/s41551-023-01045-x

Zhu, L., Yin, W., Yang, Y., Wu, F., Zeng, Z., Gu, Q., Wang, X., Zhou, C., & Ye, N. (2024). Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization. International Journal of Computer Vision 2024 132:9, 132(9), 3375-3407. https://doi.org/10.1007/S11263-024-02036-4

Downloads

Published

2026-09-11

How to Cite

Ibrahim, I., Alshar’e, M., Sanjaya, I., & Tiwari, G. (2026). Cross-Modal Representation Learning for Integrating Heterogeneous Data in AI Systems. Journal of Data Science, 2026(2), 218–235. https://doi.org/10.61453/jods.v20260213