Evaluating Data-Centric Optimization Strategies for Improving Machine Learning Generalization

Authors

  • Nia Oktaviani Universitas Bina Darma, Palembang, Indonesia
  • Elmar B. Noche Pangasinan State University – Lingayen Campus, Lingayen, Pangasinan, Philippines
  • Sonal Mohanrao Patil Bharati Vidyapeeth College of Engineering for Women, Pune, India
  • Yashomati R. Dhumal Bharati Vidyapeeth College of Engineering for Women, Pune, India

DOI:

https://doi.org/10.61453/jods.v20260212

Keywords:

Data-Centric AI, Generalization, Data Quality, Machine Learning Optimization, Model Evaluation

Abstract

Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite increasing architectural complexity. Existing studies typically evaluate data preprocessing techniques in isolation and lack a unified framework to assess their combined impact. To address this gap, this study proposes a comprehensive data-centric optimization framework that systematically integrates data cleaning, rebalancing, and targeted augmentation to enhance dataset quality. Multiple benchmark datasets with diverse characteristics were utilized, and models were trained under controlled experimental settings to isolate the effects of data refinement strategies. The results demonstrate that data-centric interventions consistently outperform model-centric improvements, with the proposed approach achieving up to 95% accuracy, significantly surpassing baseline models trained on raw data. Notably, simpler models trained on refined datasets achieve competitive or superior performance compared to more complex models, highlighting the efficiency of data-centric optimization. Statistical analysis confirms that the improvements are significant and robust across different datasets and conditions. This study aims to quantify the impact of data quality on machine learning generalization and provide empirical evidence supporting a paradigm shift toward data-centric AI. The findings offer practical guidelines for prioritizing data refinement in machine learning pipelines and contribute to the development of more robust, scalable, and efficient intelligent systems.

References

Alomar, K., Aysel, H. I., & Cai, X. (2023). Data Augmentation in Classification and Segmentation: A Survey and New Strategies. Journal of Imaging 2023, Vol. 9, 9(2). https://doi.org/10.3390/JIMAGING9020046

Altalhan, M., Algarni, A., & Turki-Hadj Alouane, M. (2025). Imbalanced Data Problem in Machine Learning: A Review. IEEE Access, 13, 13686–13699. https://doi.org/10.1109/ACCESS.2025.3531662

Asif, S., Wenhui, Y., ur-Rehman, S., ul-ain, Q., Amjad, K., Yueyang, Y., Jinhai, S., & Awais, M. (2025). Advancements and Prospects of Machine Learning in Medical Diagnostics: Unveiling the Future of Diagnostic Precision. Archives of Computational Methods in Engineering, 32(2), 853–883. https://doi.org/10.1007/s11831-024-10148-w

Bhatt, N., Bhatt, N., Prajapati, P., Sorathiya, V., Alshathri, S., & El-Shafai, W. (2024). A Data-Centric Approach to improve performance of deep learning models. Scientific Reports 2024 14:1, 14(1), 22329-. https://doi.org/10.1038/s41598-024-73643-x

Bian, K., & Priyadarshi, R. (2024). Machine Learning Optimization Techniques: A Survey, Classification, Challenges, and Future Research Issues. Archives of Computational Methods in Engineering 2024 31:7, 31(7), 4209–4233. https://doi.org/10.1007/S11831-024-10110-W

Carvalho, M., Pinho, A. J., & Brás, S. (2025). Resampling approaches to handle class imbalance: a review from a data perspective. Journal of Big Data 2025 12:1, 12(1), 71-. https://doi.org/10.1186/S40537-025-01119-4

Chicco, D., Oneto, L., & Tavazzi, E. (2022). Eleven quick tips for data cleaning and feature engineering. PLOS Computational Biology, 18(12), e1010718. https://doi.org/10.1371/JOURNAL.PCBI.1010718

Curtis, M. J., Alexander, S. P. H., Cirino, G., George, C. H., Kendall, D. A., Insel, P. A., Izzo, A. A., Ji, Y., Panettieri, R. A., Patel, H. H., Sobey, C. G., Stanford, S. C., Stanley, P., Stefanska, B., Stephens, G. J., Teixeira, M. M., Vergnolle, N., & Ahluwalia, A. (2022). Planning experiments: Updated guidance on experimental design and analysis and their reporting III. British Journal of Pharmacology, 179(15), 3907–3913. https://doi.org/10.1111/BPH.15868

Fan, C., Lei, Y., Sun, Y., Piscitelli, M. S., Chiosa, R., & Capozzoli, A. (2022). Data-centric or algorithm-centric: Exploiting the performance of transfer learning for improving building energy predictions in data-scarce context. Energy, 240, 122775. https://doi.org/10.1016/J.ENERGY.2021.122775

Foody, G. M. (2023). Challenges in the real world use of classification accuracy metrics: From recall and precision to the Matthews correlation coefficient. PLOS ONE, 18(10), e0291908. https://doi.org/10.1371/JOURNAL.PONE.0291908

Gundersen, O. E., Shamsaliei, S., & Isdahl, R. J. (2022). Do machine learning platforms provide out-of-the-box reproducibility? Future Generation Computer Systems, 126, 34–47. https://doi.org/10.1016/J.FUTURE.2021.06.014

Javed, H., El-Sappagh, S., & Abuhmed, T. (2024). Robustness in deep learning models for medical diagnostics: security and adversarial challenges towards robust AI applications. Artificial Intelligence Review 2024 58:1, 58(1), 12-. https://doi.org/10.1007/S10462-024-11005-9

Karl, F., Pielok, T., Moosbauer, J., Pfisterer, F., Coors, S., Binder, M., Schneider, L., Thomas, J., Richter, J., Lang, M., Garrido-Merchán, E. C., Branke, J., & Bischl, B. (2023). Multi-Objective Hyperparameter Optimization in Machine Learning—An Overview. ACM Transactions on Evolutionary Learning and Optimization, 3(4). https://doi.org/10.1145/3610536

Khan, K. S., Fawzy, M., & Chien, P. F. W. (2023). Integrity of randomized clinical trials: Performance of integrity tests and checklists requires assessment. International Journal of Gynecology & Obstetrics, 163(3), 733–743. https://doi.org/10.1002/IJGO.14837

Kumar, S., Datta, S., Singh, V., Singh, S. K., & Sharma, R. (2024). Opportunities and Challenges in Data-Centric AI. IEEE Access, 12, 33173–33189. https://doi.org/10.1109/ACCESS.2024.3369417

Kumaravel, A., & Vijayan, T. (2023). Comparing cost sensitive classifiers by the false-positive to false- negative ratio in diagnostic studies. Expert Systems with Applications, 227, 120303. https://doi.org/10.1016/J.ESWA.2023.120303

Ling, J., Feng, K., Wang, T., Liao, M., Yang, C., & Liu, Z. (2023). Data Modeling Techniques for Pipeline Integrity Assessment: A State-of-the-Art Survey. IEEE Transactions on Instrumentation and Measurement, 72. https://doi.org/10.1109/TIM.2023.3279910

Liu, C., Dong, Y., Xiang, W., Yang, X., Su, H., Zhu, J., Chen, Y., He, Y., Xue, H., & Zheng, S. (2024). A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking. International Journal of Computer Vision 2024 133:2, 133(2), 567–589. https://doi.org/10.1007/S11263-024-02196-3

Majeed, A., & Hwang, S. O. (2024). Towards Unlocking the Hidden Potentials of the Data-Centric AI Paradigm in the Modern Era. Applied System Innovation 2024, Vol. 7, 7(4). https://doi.org/10.3390/ASI7040054

Maleki, F., Ovens, K., Gupta, R., Reinhold, C., Spatz, A., & Forghani, R. (2022). Generalizability of Machine Learning Models: Quantitative Evaluation of Three Methodological Pitfalls. Https://Doi.Org/10.1148/Ryai.220028, 5(1). https://doi.org/10.1148/RYAI.220028

Masud, M., Eldin Rashed, A. E., & Hossain, M. S. (2020). Convolutional neural network-based models for diagnosis of breast cancer. Neural Computing and Applications 2020 34:14, 34(14), 11383–11394. https://doi.org/10.1007/S00521-020-05394-5

Mei, X., Meng, C., Liu, H., Kong, Q., Ko, T., Zhao, C., Plumbley, M. D., Zou, Y., & Wang, W. (2024). WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research. IEEE/ACM Transactions on Audio Speech and Language Processing, 32, 3339–3354. https://doi.org/10.1109/TASLP.2024.3419446

Movahedi, F., Padman, R., & Antaki, J. F. (2023). Limitations of receiver operating characteristic curve on imbalanced data: Assist device mortality risk scores. The Journal of Thoracic and Cardiovascular Surgery, 165(4), 1433-1442.e2. https://doi.org/10.1016/J.JTCVS.2021.07.041

Muhammad, L. N. (2023). Guidelines for repeated measures statistical analysis approaches with basic science research considerations. The Journal of Clinical Investigation, 133(11). https://doi.org/10.1172/JCI171058

Pan, I., Mason, L. R., & Matar, O. K. (2022). Data-centric Engineering: integrating simulation, machine learning and statistics. Challenges and opportunities. Chemical Engineering Science, 249, 117271. https://doi.org/10.1016/J.CES.2021.117271

Schwabe, D., Becker, K., Seyferth, M., Klaß, A., & Schaeffter, T. (2024). The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. Npj Digital Medicine 2024 7:1, 7(1), 203-. https://doi.org/10.1038/s41746-024-01196-4

Seedat, N., Imrie, F., & Van Der Schaar, M. (2024). Navigating Data-Centric Artificial Intelligence with DC-Check: Advances, Challenges, and Opportunities. IEEE Transactions on Artificial Intelligence, 5(6), 2589–2603. https://doi.org/10.1109/TAI.2023.3345805

Serani, A., Scholcz, T. P., & Vanzi, V. (2024). A Scoping Review on Simulation-Based Design Optimization in Marine Engineering: Trends, Best Practices, and Gaps. Archives of Computational Methods in Engineering 2024 31:8, 31(8), 4709–4737. https://doi.org/10.1007/S11831-024-10127-1

Shahbazi, N., & Asudeh, A. (2024). Reliability evaluation of individual predictions: a data-centric approach. The VLDB Journal 2024 33:4, 33(4), 1203–1230. https://doi.org/10.1007/S00778-024-00857-W

Singh, P. (2023). Systematic review of data-centric approaches in artificial intelligence and machine learning. Data Science and Management, 6(3), 144–157. https://doi.org/10.1016/J.DSM.2023.06.001

Sinha, S., & Lee, Y. M. (2024). Challenges with developing and deploying AI models and applications in industrial systems. Discover Artificial Intelligence 2024 4:1, 4(1), 55-. https://doi.org/10.1007/S44163-024-00151-2

Sun, S., Romero, A., Foehn, P., Kaufmann, E., & Scaramuzza, D. (2022). A Comparative Study of Nonlinear MPC and Differential-Flatness-Based Control for Quadrotor Agile Flight. IEEE Transactions on Robotics, 38(6), 3357–3373. https://doi.org/10.1109/TRO.2022.3177279

Talukder, M. A., Talaat, A. S., Muna, N. J., Alazab, A., Kazi, M., & Das, U. K. (2025). An explainable deep learning framework for trustworthy arrhythmia detection from ECG signals. Scientific Reports 2025 15:1, 15(1), 39496-. https://doi.org/10.1038/s41598-025-22986-0

Verkerken, M., D’hooge, L., Wauters, T., Volckaert, B., & De Turck, F. (2021). Towards Model Generalization for Intrusion Detection: Unsupervised Machine Learning Techniques. Journal of Network and Systems Management 2021 30:1, 30(1), 12-. https://doi.org/10.1007/S10922-021-09615-7

Widad, E., Saida, E., & Gahi, Y. (2023). Quality Anomaly Detection Using Predictive Techniques: An Extensive Big Data Quality Framework for Reliable Data Analysis. IEEE Access, 11, 103306–103318. https://doi.org/10.1109/ACCESS.2023.3317354

Zare, N., Macioszek, E., Granà, A., & Giuffrè, T. (2024). Blending Efficiency and Resilience in the Performance Assessment of Urban Intersections: A Novel Heuristic Informed by Literature Review. Sustainability 2024, Vol. 16, 16(6). https://doi.org/10.3390/SU16062450

Zha, D., Bhat, Z. P., Lai, K. H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2025). Data-centric Artificial Intelligence: A Survey. ACM Computing Surveys, 57(5), 129. https://doi.org/10.1145/3711118

Zhu, J. J., Yang, M., & Ren, Z. J. (2023). Machine Learning in Environmental Research: Common Pitfalls and Best Practices. Environmental Science & Technology, 57(46), 17671–17689. https://doi.org/10.1021/ACS.EST.3C00026

Downloads

Published

2026-09-11

How to Cite

Oktaviani, N., Noche, E. B., Patil, S. M., & Dhumal, Y. R. (2026). Evaluating Data-Centric Optimization Strategies for Improving Machine Learning Generalization. Journal of Data Science, 2026(2), 200–217. https://doi.org/10.61453/jods.v20260212