Ariyadi, Tamsir and Noche, Elmar B. and Pandey, Nisha and Sinha, Vivek Kumar (2026) Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions. Journal of Data Science, 2026 (10). pp. 168-182. ISSN 2805-5160
|
Text
jods2026_10.pdf - Published Version Available under License Creative Commons Attribution. Download (193kB) |
|
|
Text
928 - Published Version Available under License Creative Commons Attribution. Download (39kB) |
Abstract
Extreme class imbalance remains a persistent challenge in machine learning, particularly in high-impact domains such as fraud detection, medical diagnosis, and risk analysis, where minority classes represent critical outcomes. Conventional models often fail in such settings due to their bias toward majority classes, resulting in poor minority detection despite high overall accuracy. Although various data-level and algorithm-level techniques have been proposed, existing studies typically evaluate them in isolation and lack a comprehensive understanding of their effectiveness across different imbalance conditions. To address this gap, this study proposes a systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels. Multiple benchmark datasets are utilized, and experiments are conducted using standardized preprocessing, controlled imbalance simulation, and repeated trials to ensure robustness. Performance is assessed using imbalance-aware metrics, including precision, recall, F1-score, and ROC-AUC. The results indicate that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced performance across all scenarios. In particular, hybrid models demonstrate superior minority class recall and F1-score while maintaining competitive precision, especially under extreme imbalance conditions. The primary goal of this research is to provide a comprehensive evaluation of imbalance-handling strategies and offer practical guidance for selecting appropriate techniques based on dataset characteristics. The findings highlight the importance of combining data-centric and model-centric approaches to enhance robustness and reliability in imbalanced learning environments. The results demonstrate up to a 32% improvement in recall compared to baseline models.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Class Imbalance, Imbalanced Learning, Cost-Sensitive Learning, Data Preprocessing, Model Evaluation |
| Subjects: | Q Science > QA Mathematics Q Science > QA Mathematics > QA75 Electronic computers. Computer science Q Science > QA Mathematics > QA76 Computer software |
| Depositing User: | Unnamed user with email masilah.mansor@newinti.edu.my |
| Date Deposited: | 10 Sep 2026 08:54 |
| Last Modified: | 10 Sep 2026 08:54 |
| URI: | http://eprints.intimal.edu.my/id/eprint/2365 |
Actions (login required)
![]() |
View Item |
