Evaluating Data-Centric Optimization Strategies for Improving Machine Learning Generalization
DOI:
https://doi.org/10.61453/jods.v20260212Keywords:
Data-Centric AI, Generalization, Data Quality, Machine Learning Optimization, Model EvaluationAbstract
Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite increasing architectural complexity. Existing studies typically evaluate data preprocessing techniques in isolation and lack a unified framework to assess their combined impact. To address this gap, this study proposes a comprehensive data-centric optimization framework that systematically integrates data cleaning, rebalancing, and targeted augmentation to enhance dataset quality. Multiple benchmark datasets with diverse characteristics were utilized, and models were trained under controlled experimental settings to isolate the effects of data refinement strategies. The results demonstrate that data-centric interventions consistently outperform model-centric improvements, with the proposed approach achieving up to 95% accuracy, significantly surpassing baseline models trained on raw data. Notably, simpler models trained on refined datasets achieve competitive or superior performance compared to more complex models, highlighting the efficiency of data-centric optimization. Statistical analysis confirms that the improvements are significant and robust across different datasets and conditions. This study aims to quantify the impact of data quality on machine learning generalization and provide empirical evidence supporting a paradigm shift toward data-centric AI. The findings offer practical guidelines for prioritizing data refinement in machine learning pipelines and contribute to the development of more robust, scalable, and efficient intelligent systems.
References
Alomar, K., Aysel, H. I., & Cai, X. (2023). Data Augmentation in Classification and Segmentation: A Survey and New Strategies. Journal of Imaging 2023, Vol. 9, 9(2). https://doi.org/10.3390/JIMAGING9020046
Altalhan, M., Algarni, A., & Turki-Hadj Alouane, M. (2025). Imbalanced Data Problem in Machine Learning: A Review. IEEE Access, 13, 13686–13699. https://doi.org/10.1109/ACCESS.2025.3531662
Asif, S., Wenhui, Y., ur-Rehman, S., ul-ain, Q., Amjad, K., Yueyang, Y., Jinhai, S., & Awais, M. (2025). Advancements and Prospects of Machine Learning in Medical Diagnostics: Unveiling the Future of Diagnostic Precision. Archives of Computational Methods in Engineering, 32(2), 853–883. https://doi.org/10.1007/s11831-024-10148-w
Bhatt, N., Bhatt, N., Prajapati, P., Sorathiya, V., Alshathri, S., & El-Shafai, W. (2024). A Data-Centric Approach to improve performance of deep learning models. Scientific Reports 2024 14:1, 14(1), 22329-. https://doi.org/10.1038/s41598-024-73643-x
Bian, K., & Priyadarshi, R. (2024). Machine Learning Optimization Techniques: A Survey, Classification, Challenges, and Future Research Issues. Archives of Computational Methods in Engineering 2024 31:7, 31(7), 4209–4233. https://doi.org/10.1007/S11831-024-10110-W
Carvalho, M., Pinho, A. J., & Brás, S. (2025). Resampling approaches to handle class imbalance: a review from a data perspective. Journal of Big Data 2025 12:1, 12(1), 71-. https://doi.org/10.1186/S40537-025-01119-4
Chicco, D., Oneto, L., & Tavazzi, E. (2022). Eleven quick tips for data cleaning and feature engineering. PLOS Computational Biology, 18(12), e1010718. https://doi.org/10.1371/JOURNAL.PCBI.1010718
Curtis, M. J., Alexander, S. P. H., Cirino, G., George, C. H., Kendall, D. A., Insel, P. A., Izzo, A. A., Ji, Y., Panettieri, R. A., Patel, H. H., Sobey, C. G., Stanford, S. C., Stanley, P., Stefanska, B., Stephens, G. J., Teixeira, M. M., Vergnolle, N., & Ahluwalia, A. (2022). Planning experiments: Updated guidance on experimental design and analysis and their reporting III. British Journal of Pharmacology, 179(15), 3907–3913. https://doi.org/10.1111/BPH.15868
Fan, C., Lei, Y., Sun, Y., Piscitelli, M. S., Chiosa, R., & Capozzoli, A. (2022). Data-centric or algorithm-centric: Exploiting the performance of transfer learning for improving building energy predictions in data-scarce context. Energy, 240, 122775. https://doi.org/10.1016/J.ENERGY.2021.122775
Foody, G. M. (2023). Challenges in the real world use of classification accuracy metrics: From recall and precision to the Matthews correlation coefficient. PLOS ONE, 18(10), e0291908. https://doi.org/10.1371/JOURNAL.PONE.0291908
Gundersen, O. E., Shamsaliei, S., & Isdahl, R. J. (2022). Do machine learning platforms provide out-of-the-box reproducibility? Future Generation Computer Systems, 126, 34–47. https://doi.org/10.1016/J.FUTURE.2021.06.014
Javed, H., El-Sappagh, S., & Abuhmed, T. (2024). Robustness in deep learning models for medical diagnostics: security and adversarial challenges towards robust AI applications. Artificial Intelligence Review 2024 58:1, 58(1), 12-. https://doi.org/10.1007/S10462-024-11005-9
Karl, F., Pielok, T., Moosbauer, J., Pfisterer, F., Coors, S., Binder, M., Schneider, L., Thomas, J., Richter, J., Lang, M., Garrido-Merchán, E. C., Branke, J., & Bischl, B. (2023). Multi-Objective Hyperparameter Optimization in Machine Learning—An Overview. ACM Transactions on Evolutionary Learning and Optimization, 3(4). https://doi.org/10.1145/3610536
Khan, K. S., Fawzy, M., & Chien, P. F. W. (2023). Integrity of randomized clinical trials: Performance of integrity tests and checklists requires assessment. International Journal of Gynecology & Obstetrics, 163(3), 733–743. https://doi.org/10.1002/IJGO.14837
Kumar, S., Datta, S., Singh, V., Singh, S. K., & Sharma, R. (2024). Opportunities and Challenges in Data-Centric AI. IEEE Access, 12, 33173–33189. https://doi.org/10.1109/ACCESS.2024.3369417
Kumaravel, A., & Vijayan, T. (2023). Comparing cost sensitive classifiers by the false-positive to false- negative ratio in diagnostic studies. Expert Systems with Applications, 227, 120303. https://doi.org/10.1016/J.ESWA.2023.120303
Ling, J., Feng, K., Wang, T., Liao, M., Yang, C., & Liu, Z. (2023). Data Modeling Techniques for Pipeline Integrity Assessment: A State-of-the-Art Survey. IEEE Transactions on Instrumentation and Measurement, 72. https://doi.org/10.1109/TIM.2023.3279910
Liu, C., Dong, Y., Xiang, W., Yang, X., Su, H., Zhu, J., Chen, Y., He, Y., Xue, H., & Zheng, S. (2024). A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking. International Journal of Computer Vision 2024 133:2, 133(2), 567–589. https://doi.org/10.1007/S11263-024-02196-3
Majeed, A., & Hwang, S. O. (2024). Towards Unlocking the Hidden Potentials of the Data-Centric AI Paradigm in the Modern Era. Applied System Innovation 2024, Vol. 7, 7(4). https://doi.org/10.3390/ASI7040054
Maleki, F., Ovens, K., Gupta, R., Reinhold, C., Spatz, A., & Forghani, R. (2022). Generalizability of Machine Learning Models: Quantitative Evaluation of Three Methodological Pitfalls. Https://Doi.Org/10.1148/Ryai.220028, 5(1). https://doi.org/10.1148/RYAI.220028
Masud, M., Eldin Rashed, A. E., & Hossain, M. S. (2020). Convolutional neural network-based models for diagnosis of breast cancer. Neural Computing and Applications 2020 34:14, 34(14), 11383–11394. https://doi.org/10.1007/S00521-020-05394-5
Mei, X., Meng, C., Liu, H., Kong, Q., Ko, T., Zhao, C., Plumbley, M. D., Zou, Y., & Wang, W. (2024). WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research. IEEE/ACM Transactions on Audio Speech and Language Processing, 32, 3339–3354. https://doi.org/10.1109/TASLP.2024.3419446
Movahedi, F., Padman, R., & Antaki, J. F. (2023). Limitations of receiver operating characteristic curve on imbalanced data: Assist device mortality risk scores. The Journal of Thoracic and Cardiovascular Surgery, 165(4), 1433-1442.e2. https://doi.org/10.1016/J.JTCVS.2021.07.041
Muhammad, L. N. (2023). Guidelines for repeated measures statistical analysis approaches with basic science research considerations. The Journal of Clinical Investigation, 133(11). https://doi.org/10.1172/JCI171058
Pan, I., Mason, L. R., & Matar, O. K. (2022). Data-centric Engineering: integrating simulation, machine learning and statistics. Challenges and opportunities. Chemical Engineering Science, 249, 117271. https://doi.org/10.1016/J.CES.2021.117271
Schwabe, D., Becker, K., Seyferth, M., Klaß, A., & Schaeffter, T. (2024). The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. Npj Digital Medicine 2024 7:1, 7(1), 203-. https://doi.org/10.1038/s41746-024-01196-4
Seedat, N., Imrie, F., & Van Der Schaar, M. (2024). Navigating Data-Centric Artificial Intelligence with DC-Check: Advances, Challenges, and Opportunities. IEEE Transactions on Artificial Intelligence, 5(6), 2589–2603. https://doi.org/10.1109/TAI.2023.3345805
Serani, A., Scholcz, T. P., & Vanzi, V. (2024). A Scoping Review on Simulation-Based Design Optimization in Marine Engineering: Trends, Best Practices, and Gaps. Archives of Computational Methods in Engineering 2024 31:8, 31(8), 4709–4737. https://doi.org/10.1007/S11831-024-10127-1
Shahbazi, N., & Asudeh, A. (2024). Reliability evaluation of individual predictions: a data-centric approach. The VLDB Journal 2024 33:4, 33(4), 1203–1230. https://doi.org/10.1007/S00778-024-00857-W
Singh, P. (2023). Systematic review of data-centric approaches in artificial intelligence and machine learning. Data Science and Management, 6(3), 144–157. https://doi.org/10.1016/J.DSM.2023.06.001
Sinha, S., & Lee, Y. M. (2024). Challenges with developing and deploying AI models and applications in industrial systems. Discover Artificial Intelligence 2024 4:1, 4(1), 55-. https://doi.org/10.1007/S44163-024-00151-2
Sun, S., Romero, A., Foehn, P., Kaufmann, E., & Scaramuzza, D. (2022). A Comparative Study of Nonlinear MPC and Differential-Flatness-Based Control for Quadrotor Agile Flight. IEEE Transactions on Robotics, 38(6), 3357–3373. https://doi.org/10.1109/TRO.2022.3177279
Talukder, M. A., Talaat, A. S., Muna, N. J., Alazab, A., Kazi, M., & Das, U. K. (2025). An explainable deep learning framework for trustworthy arrhythmia detection from ECG signals. Scientific Reports 2025 15:1, 15(1), 39496-. https://doi.org/10.1038/s41598-025-22986-0
Verkerken, M., D’hooge, L., Wauters, T., Volckaert, B., & De Turck, F. (2021). Towards Model Generalization for Intrusion Detection: Unsupervised Machine Learning Techniques. Journal of Network and Systems Management 2021 30:1, 30(1), 12-. https://doi.org/10.1007/S10922-021-09615-7
Widad, E., Saida, E., & Gahi, Y. (2023). Quality Anomaly Detection Using Predictive Techniques: An Extensive Big Data Quality Framework for Reliable Data Analysis. IEEE Access, 11, 103306–103318. https://doi.org/10.1109/ACCESS.2023.3317354
Zare, N., Macioszek, E., Granà, A., & Giuffrè, T. (2024). Blending Efficiency and Resilience in the Performance Assessment of Urban Intersections: A Novel Heuristic Informed by Literature Review. Sustainability 2024, Vol. 16, 16(6). https://doi.org/10.3390/SU16062450
Zha, D., Bhat, Z. P., Lai, K. H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2025). Data-centric Artificial Intelligence: A Survey. ACM Computing Surveys, 57(5), 129. https://doi.org/10.1145/3711118
Zhu, J. J., Yang, M., & Ren, Z. J. (2023). Machine Learning in Environmental Research: Common Pitfalls and Best Practices. Environmental Science & Technology, 57(46), 17671–17689. https://doi.org/10.1021/ACS.EST.3C00026
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.