Robust Evaluation Metrics for Assessing Machine Learning Performance Beyond Accuracy
DOI:
https://doi.org/10.61453/jods.v20260216Keywords:
Machine Learning Evaluation, Performance Metrics, Robustness Assessment, Calibration Metrics, Trustworthy Artificial IntelligenceAbstract
The widespread deployment of machine learning (ML) systems in critical domains has exposed the limitations of accuracy-centric evaluation, particularly under conditions involving class imbalance, noise, and distributional shifts. Existing studies frequently employ alternative metrics in isolation and lack a unified framework capable of systematically assessing model robustness, reliability, and decision sensitivity across varying data conditions. To address this gap, this study proposes a structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment. A quantitative experimental design is employed using multiple benchmark datasets with varying statistical characteristics, including balanced and imbalanced distributions. Controlled perturbation scenarios—including noise injection, class imbalance manipulation, and distribution shift simulation—are introduced to emulate realistic deployment environments. Several machine learning models, namely Logistic Regression, Support Vector Machines, Random Forest, and Multi-Layer Perceptron (MLP), are evaluated using metrics such as Accuracy, F1-score, ROC-AUC, PR-AUC, Brier Score, and Expected Calibration Error (ECE). The experimental results demonstrate that accuracy consistently overestimates model effectiveness under adverse conditions, while alternative metrics reveal substantial hidden weaknesses in minority class detection and probability reliability. Among the evaluated models, MLP achieved the strongest overall performance, obtaining a ROC-AUC of 0.94 and PR-AUC of 0.89 under baseline conditions. Furthermore, calibration-oriented metrics exhibited significantly higher sensitivity to perturbation severity compared to accuracy. This study contributes to the advancement of trustworthy artificial intelligence by promoting a comprehensive, context-aware, and robustness-oriented evaluation framework capable of supporting more reliable real-world ML deployment.
References
Abedin, T., Xu, H., & Uddin, S. (2026). The impact of K selection in K fold cross-validation on bias and variance in supervised learning models. Scientific Reports 2026 16:1, 16(1), 6084 https://doi.org/10.1038/s41598-026-37247-x
Aguiar, G., Krawczyk, B., & Cano, A. (2024). A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework. Machine Learning, 113(7), 4165-4243. https://doi.org/10.1007/s10994-023-06353-6
Ahin, E. S.,, Naciye,, Arslan, N., Durmus, o zdemir,, Durmus,o, D., & Durmus, o zdemir, D. (2024). Unlocking the black box: an in-depth review on interpretability, explainability, and reliability in deep learning. Neural Computing and Applications 2024 37:2, 37(2), 859– 965. https://doi.org/10.1007/S00521-024-10437-2
Bagherian, A., Chauhan, G., & Srivastav, A. L. (2023). Data-driven prioritization of performance variables for flexible manufacturing systems: revealing key metrics with the best-worst method. The International Journal of Advanced Manufacturing Technology 2023 130:5, 130(5), 3081-3102. https://doi.org/10.1007/S00170-023-12784-1
Bender, A., Schneider, N., Segler, M., Patrick Walters, W., Engkvist, O., & Rodrigues, T. (2022). Evaluation guidelines for machine learning tools in the chemical sciences. Nature Reviews Chemistry 2022 6:6, 6(6), 428–442. https://doi.org/10.1038/s41570-022-00391-9
Cascone, L., Nappi, M., Pero, C., & Wang, X. (2026). A framework for bias-aware dataset evaluation in soft facial attribute recognition. Pattern Recognition, 172, 112416. https://doi.org/10.1016/J.PATCOG.2025.112416
Cheung, G. W., Cooper-Thomas, H. D., Lau, R. S., & Wang, L. C. (2023). Reporting reliability, convergent and discriminant validity with structural equation modeling: A review and best-practice recommendations. Asia Pacific Journal of Management 2023 41:2, 41(2), 745-783. https://doi.org/10.1007/S10490-023-09871-Y
Imani, M., Joudaki, M., Bagheri, A., & Arabnia, H. R. (2026). Why ROC-AUC Is Misleading for Highly Imbalanced Data: In-Depth Evaluation of MCC, F2-Score, H-Measure, and AUC-Based Metrics Across Diverse Classifiers. Technologies 2026, Vol. 14, 14(1), 54. https://doi.org/10.3390/TECHNOLOGIES14010054
Li, W., & Chai, Y. (2022). Assessing and Enhancing Adversarial Robustness of Predictive Analytics: An Empirically Tested Design Framework. Journal of Management Information Systems, 39(2), 542-572. https://doi.org/10.1080/07421222.2022.2063549
Lu, H. S., & Daugherty, A. (2022). Key Factors for Improving Rigor and Reproducibility: Guidelines, Peer Reviews, and Journal Technical Reviews. Frontiers in Cardiovascular Medicine, 9, 856102. https://doi.org/10.3389/FCVM.2022.856102
Maleki, F., Ovens, K., Gupta, R., Reinhold, C., Spatz, A., & Forghani, R. (2022). Generalizability of Machine Learning Models: Quantitative Evaluation of Three Methodological Pitfalls. Radiology: Artificial Intelligence. 5(1). https://doi.org/10.1148/RYAI.220028
Mamalakis, M., Banerjee, A., Ray, S., Wilkie, C., Clayton, R. H., Swift, A. J., Panoutsos, G., & Vorselaars, B. (2024). Deep multi-metric training: the need of multi-metric curve evaluation to avoid weak learning. Neural Computing and Applications 2024 36:30, 36(30), 18841-18862. https://doi.org/10.1007/S00521-024-10182-6
Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports 2024 14:1, 14(1), 6086-. https://doi.org/10.1038/s41598-024-56706-x
Snyder, H. (2024). Designing the literature review for a strong contribution. Journal of Decision Systems, 33(4), 551-558. https://doi.org/10.1080/12460125.2023.2197704
Song, Y., Wang, T., Cai, P., Mondal, S. K., & Sahoo, J. P. (2023). A Comprehensive Survey of Few-shot Learning: Evolution, Applications, Challenges, and Opportunities. ACM Computing 55(13s). Surveys, https://doi.org/10.1145/3582688
Strielkowski, W., Vlasov, A., Selivanov, K., Muraviev, K., & Shakhnov, V. (2023). Prospects and A Challenges of the Machine Learning and Data-Driven Methods for the Predictive Analysis of Power Systems: Review. Energies 2023, Vol. 16, 16(10). https://doi.org/10.3390/EN16104025
Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, fl-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data 2025 12:1, 12(1), 268-. https://doi.org/10.1186/S40537-025-01313-4
Tamak, S., Eslami, Y., & Cunha, C. Da. (2025). Validation of multidimensional performance assessment models using hierarchical clustering. Expert Systems with Applications, 290, 128446. https://doi.org/10.1016/J.ESWA.2025.128446
Tran, A. T., Zeevi, T., & Payabvash, S. (2025). Strategies to Improve the Robustness and Generalizability of Deep Learning Segmentation and Classification in Neuroimaging. 5, BioMedInformatics 2025, Vol. 5(2). https://doi.org/10.3390/BIOMEDINFORMATICS5020020
Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and Al are useful: A systematic review and meta-analysis. Nature Human Behaviour 2024 8:12, 8(12), 2293-2303. https://doi.org/10.1038/s41562-024-02024-1
Wang, A. X., Chukova, S. S., Simpson, C. R., & Nguyen, B. P. (2024). Challenges and opportunities of generative models on tabular data. Applied Soft Computing, 166, 112223. https://doi.org/10.1016/J.ASOC.2024.112223
Wang, H., Liang, Q., Hancock, J. T., & Khoshgoftaar, T. M. (2024). Feature selection strategies: a comparative analysis of SHAP-value and importance-based methods. Journal of Big Data 2024 11:1, 11(1), 44-. https://doi.org/10.1186/S40537-024-00905-W
Wei, Z., Wang, Y., Gao, Y., Wang, S., Li, P., Si, D., Gao, Y., Wu, S., Li, D., Dong, K., Yang, Benchmarking algorithms for generalizable single-cell perturbation response prediction. Nature Methods 2025 23:2, 23(2), 451-464. https://doi.org/10.1038/s41592-025-02980-0
Zhu, J. J., Yang, M., & Ren, Z. J. (2023). Machine Learning in Environmental Research: Common Pitfalls and Best Practices. Environmental Science & Technology, 57(46), 17671-17689. https://doi.org/10.1021/ACS.EST.3C00026
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.