Human-in-the-Loop Data Science: Enhancing Model Performance Through Interactive Learning Mechanisms
DOI:
https://doi.org/10.61453/jods.v20260211Keywords:
Human-In-The-Loop, Interactive Learning, Collaborative Intelligence, Data Science Workflows, Model RefinementAbstract
Purely automated machine learning systems often struggle to incorporate domain knowledge and contextual reasoning, resulting in reduced performance when handling ambiguous, noisy, or complex real-world data. Although existing approaches such as active learning and semi-supervised learning partially address these limitations, they typically treat human input as an auxiliary component rather than an integral part of the learning process. This creates a critical research gap in developing frameworks that systematically and iteratively integrate human expertise into model training. To address this issue, this study proposes a human-in-the-loop (HITL) data science framework that embeds structured human feedback into the machine learning lifecycle, enabling continuous refinement of model predictions through interactive learning mechanisms. The proposed framework is evaluated using real-world datasets requiring expert judgment, with experiments designed to compare fully automated models, limited human interaction, and continuous HITL integration. The results demonstrate that the proposed approach achieves superior performance, including improved predictive accuracy, faster convergence, and significant reduction in labeling errors. Notably, the HITL framework shows consistent robustness under noisy and ambiguous data conditions, outperforming baseline models while maintaining stable learning behavior. This research aims to enhance the reliability, interpretability, and adaptability of machine learning systems by leveraging collaborative intelligence between humans and machines. The findings highlight that structured human feedback can serve as an effective learning signal, enabling more accurate and trustworthy data-driven models.
References
Adnan, H. S., Shidani, A., Clifton, L., Bankhead, C. R., & Perera-Salazar, R. (2025). Implementation framework for AI deployment at scale in healthcare systems. IScience, 28(5), 112406. https://doi.org/10.1016/j.isci.2025.112406
Arunthavanathan, R., Sajid, Z., Khan, F., & Pistikopoulos, E. (2024). Artificial intelligence – Human intelligence conflict and its impact on process system safety. Digital Chemical Engineering, 11, 100151. https://doi.org/10.1016/J.DCHE.2024.100151
Chen, H., Li, S., Fan, J., Duan, A., Yang, C., Navarro-Alarcon, D., & Zheng, P. (2025). Human-in-the-Loop Robot Learning for Smart Manufacturing: A Human-Centric Perspective. IEEE Transactions on Automation Science and Engineering, 22, 11062–11086. https://doi.org/10.1109/TASE.2025.3528051
Chen, V., Bhatt, U., Heidari, H., Weller, A., & Talwalkar, A. (2023). Perspectives on incorporating expert feedback into model updates. Patterns, 4(7), 100780. https://doi.org/10.1016/J.PATTER.2023.100780
Cowley, H. P., Natter, M., Gray-Roncal, K., Rhodes, R. E., Johnson, E. C., Drenkow, N., Shead, T. M., Chance, F. S., Wester, B., & Gray-Roncal, W. (2022). A framework for rigorous evaluation of human performance in human and machine learning comparison studies. Scientific Reports 2022 12:1, 12(1), 5444-. https://doi.org/10.1038/s41598-022-08078-3
Da Bormida, M., Cugurra, I., Koussouris, S., Crippa, D., Hans, C., Hellbach, R., Tabone, A., Bibikas, D., In Cugurra, M. D. B., Koussouris, S., Crippa, D., Hans, C., Hellbach Biba-Bremer, R., Tabone, A., & Bibikas, D. (2026). The Cross-Fertilization Between the Human-in-the-Loop Approach and the Explainable AI Techniques Toward Trustworthiness. Artificial Intelligence, Data and Robotics, 77–104. https://doi.org/10.1007/978-3-032-10561-5_5
El-Sayed, A., Nasr, A., Mohamed, Y., Alaaeldin, A., Ali, M., Salah, O., Khalid, A., & Lazem, S. (2025). A data centric HitL framework for conducting a systematic error analysis of NLP datasets using explainable AI. Scientific Reports 2025 15:1, 15(1), 30406-. https://doi.org/10.1038/s41598-025-13452-y
Fernandes, P., Madaan, A., Liu, E., Farinhas, A., Martins, P. H., Bertsch, A., de Souza, J. G. C., Zhou, S., Wu, T., Neubig, G., & Martins, A. F. T. (2023). Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation. Transactions of the Association for Computational Linguistics, 11, 1643–1668. https://doi.org/10.1162/TACL_A_00626
Fischer, F., Fleig, A., Klar, M., & Müller, J. (2022). Optimal Feedback Control for Modeling Human-Computer Interaction. ACM Transactions on Computer-Human Interaction, 29(6). https://doi.org/10.1145/3524122
Gómez-Carmona, O., Casado-Mansilla, D., López-de-Ipiña, D., & García-Zubia, J. (2024). Human-in-the-loop machine learning: Reconceptualizing the role of the user in interactive approaches. Internet of Things, 25, 101048. https://doi.org/10.1016/J.IOT.2023.101048
Gong, J. (2024). Pushing the Boundary: Specialising Deep Configuration Performance Learning. https://arxiv.org/pdf/2407.02706
Hong, S., Yu, W., & Chai, T. (2025). A Human-in-the-Loop Framework for Interactive and Explainable Data-Driven Modeling. IEEE Transactions on Emerging Topics in Computational Intelligence. https://doi.org/10.1109/TETCI.2025.3637807
Javed, H., El-Sappagh, S., & Abuhmed, T. (2024). Robustness in deep learning models for medical diagnostics: security and adversarial challenges towards robust AI applications. Artificial Intelligence Review 2024 58:1, 58(1), 12-. https://doi.org/10.1007/S10462-024-11005-9
Jin, C., Xiao, Y., Wu, H., Ji, X., Li, G., Shuai, J., Yang, P., & Xiong, L. (2025). Human-model interaction-based decision support system for optimizing food safety assessment. Food Research International, 208, 116156. https://doi.org/10.1016/J.FOODRES.2025.116156
Liu, F., & Demosthenes, P. (2022). Real-world data: a brief review of the methods, applications, challenges and opportunities. BMC Medical Research Methodology 2022 22:1, 22(1), 287-. https://doi.org/10.1186/S12874-022-01768-6
Maier, U., & Klotz, C. (2022). Personalized feedback in digital learning environments: Classification framework and literature review. Computers and Education: Artificial Intelligence, 3, 100080. https://doi.org/10.1016/J.CAEAI.2022.100080
Mehta, S. A., Meng, F., Bajcsy, A., & Losey, D. P. (2024). StROL: Stabilized and Robust Online Learning from Humans. IEEE Robotics and Automation Letters, 9(3), 2303–2310. https://doi.org/10.1109/LRA.2024.3354626
Menghani, G. (2023). Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better. ACM Computing Surveys, 55(12). https://doi.org/10.1145/3578938
Messer, M., Brown, N. C. C., Kölling, M., & Shi, M. (2025). How Consistent Are Humans When Grading Programming Assignments? ACM Transactions on Computing Education, 25(4), 49. https://doi.org/10.1145/3759256
Mosqueira-Rey, E., Hernández-Pereira, E., Alonso-Ríos, D., Bobes-Bascarán, J., & Fernández-Leal, Á. (2022). Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review 2022 56:4, 56(4), 3005–3054. https://doi.org/10.1007/S10462-022-10246-W
Mosqueira-Rey, E., Hernández-Pereira, E., Bobes-Bascarán, J., Alonso-Ríos, D., Pérez-Sánchez, A., Fernández-Leal, Á., Moret-Bonillo, V., Vidal-Ínsua, Y., & Vázquez-Rivera, F. (2023). Addressing the data bottleneck in medical deep learning models using a human-in-the-loop machine learning approach. Neural Computing and Applications 2023 36:5, 36(5), 2597–2616. https://doi.org/10.1007/S00521-023-09197-2
Mwamsojo, N., Lehmann, F., Merghem, K., Frignac, Y., & Benkelfat, B. E. (2024). A stochastic optimization technique for hyperparameter tuning in reservoir computing. Neurocomputing, 574, 127262. https://doi.org/10.1016/J.NEUCOM.2024.127262
Myren, S., Parikh, N., Rael, R., Flynn, G., Higdon, D., & Casleton, E. (2026). Evaluation of Seismic Artificial Intelligence with Uncertainty. Seismological Research Letters, 97(1), 471–486. https://doi.org/10.1785/0220240444
Nadimi-Shahraki, M. H., Zamani, H., Asghari Varzaneh, Z., & Mirjalili, S. (2023). A Systematic Review of the Whale Optimization Algorithm: Theoretical Foundation, Improvements, and Hybridizations. Archives of Computational Methods in Engineering 2023 30:7, 30(7), 4113–4159. https://doi.org/10.1007/S11831-023-09928-7
Pan, S., Yang, B., Wang, S., Guo, Z., Wang, L., Liu, J., & Wu, S. (2023). Oil well production prediction based on CNN-LSTM model with self-attention mechanism. Energy, 284, 128701. https://doi.org/10.1016/J.ENERGY.2023.128701
Pareschi, R. (2024). Beyond Human and Machine: An Architecture and Methodology Guideline for Centaurian Design. Sci 2024, Vol. 6, 6(4). https://doi.org/10.3390/SCI6040071
Ren, M., Chen, N., & Qiu, H. (2023). Human-machine Collaborative Decision-making: An Evolutionary Roadmap Based on Cognitive Intelligence. International Journal of Social Robotics 2023 15:7, 15(7), 1101–1114. https://doi.org/10.1007/S12369-023-01020-1
Retzlaff, C. O., Das, S., Wayllace, C., Mousavi, P., Afshari, M., Yang, T., Saranti, A., Angerschmid, A., Taylor, M. E., & Holzinger, A. (2024). Human-in-the-Loop Reinforcement Learning: A Survey and Position on Requirements, Challenges, and Opportunities. Journal of Artificial Intelligence Research, 79, 359–415. https://doi.org/10.1613/JAIR.1.15348
Sarker, I. H. (2022). Machine Learning for Intelligent Data Analysis and Automation in Cybersecurity: Current and Future Prospects. Annals of Data Science 2022 10:6, 10(6), 1473–1498. https://doi.org/10.1007/S40745-022-00444-2
Shadiev, R., & Feng, Y. (2024). Using automated corrective feedback tools in language learning: a review study. Interactive Learning Environments, 32(6), 2538–2566. https://doi.org/10.1080/10494820.2022.2153145
Shaheen, A., Khaldy, M. Al, Alzyadat, W., & Alhroob, A. (2025). AI-Driven Augmented Software Engineering: Leveraging Cognitive Models for Enhanced Code Generation. Journal of Computational and Cognitive Engineering, 2025(00), 1–9. https://doi.org/10.47852/BONVIEWJCCE52026123
Silva Filho, T., Song, H., Perello-Nieto, M., Santos-Rodriguez, R., Kull, M., & Flach, P. (2023). Classifier calibration: a survey on how to assess and improve predicted class probabilities. Machine Learning 2023 112:9, 112(9), 3211–3260. https://doi.org/10.1007/S10994-023-06336-7
Tung Khuat, T., Jacob Kedziora, D., & Gabrys, B. (2023). The Roles and Modes of Human Interactions with Automated Machine Learning Systems: A Critical Review and Perspectives. Foundations and Trends in Human-Computer Interaction, 17(3–4), 195–387. https://doi.org/10.1561/1100000091
Yang, X., Song, Z., King, I., & Xu, Z. (2023). A Survey on Deep Semi-Supervised Learning. IEEE Transactions on Knowledge and Data Engineering, 35(9), 8934–8954. https://doi.ieeecomputersociety.org/10.1109/TKDE.2022.3220219
Yu, Q., Teixeira, A. P., Liu, K., & Guedes Soares, C. (2022). Framework and application of multi-criteria ship collision risk assessment. Ocean Engineering, 250, 111006. https://doi.org/10.1016/J.OCEANENG.2022.111006
Zhou, R., Zhang, Z., Fu, H., Zhang, L., Li, L., Huang, G., Li, F., Yang, X., Dong, Y., Zhang, Y. T., & Liang, Z. (2024). PR-PL: A Novel Prototypical Representation Based Pairwise Learning Framework for Emotion Recognition Using EEG Signals. IEEE Transactions on Affective Computing, 15(2), 657–670. https://doi.org/10.1109/TAFFC.2023.3288118
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.