Evaluating the Impact of SMOTE-Based Data Balancing on Decision Bias and Algorithmic Fairness in XGBoost-Based Student Dropout Prediction

Les Endahti, Taqwa Hariguna, Dhanar Intan Surya Saputra

Abstract


Student dropout prediction is a key application of Educational Data Mining for supporting early intervention in higher education. However, previous studies have primarily focused on improving predictive accuracy, while the effects of data balancing on decision bias and algorithmic fairness remain underexplored. This study proposes a comprehensive evaluation framework that integrates predictive performance, decision bias, and algorithmic fairness to assess the impact of the Synthetic Minority Over-sampling Technique (SMOTE) on Extreme Gradient Boosting (XGBoost) for student dropout prediction. Experiments were conducted using the publicly available Predict Students Dropout and Academic Success dataset containing 4,424 student records. After excluding the Enrolled class, the dataset was transformed into a binary classification problem consisting of 2,209 Graduate (60.9%) and 1,421 Dropout (39.1%) instances. Two models were compared: a baseline XGBoost classifier and an XGBoost classifier trained with SMOTE. Predictive performance was evaluated using Accuracy, Precision, Recall, F1-score, and ROC-AUC, while decision bias and algorithmic fairness were assessed using the False Negative Rate (FNR), Statistical Parity Difference (SPD), Disparate Impact (DI), Equal Opportunity Difference (EOD), and Average Odds Difference (AOD). The baseline model achieved higher Accuracy (93.11% vs. 92.29%), Precision (91.49% vs. 89.86%), F1-score (91.17% vs. 90.18%), and a lower FNR (0.0915 vs. 0.0951), whereas both models produced comparable ROC-AUC values (0.972). McNemar's test indicated that the difference in predictive performance was not statistically significant (p = 0.264). Although SMOTE did not improve predictive performance, it produced modest reductions in Statistical Parity Difference (0.2310–0.2218), Equal Opportunity Difference (0.0298–0.0233), and Average Odds Difference (0.0295–0.0235), indicating a slight improvement in fairness metrics while maintaining comparable discrimination capability. These findings highlight the trade-off between predictive performance and algorithmic fairness and demonstrate that evaluating predictive performance together with decision bias and fairness provides a more comprehensive assessment of educational machine learning models, supporting the development of responsible AI-based educational decision-support systems.

Keywords


Student Dropout Prediction;Educational Data Mining;XGBoost;SMOTE;Decision Bias;Algorithmic Fairness

Full Text:

Link Download

References


Amin, A., Anwar, S., Adnan, A., Nawaz, M., Howard, N., Qadir, J., Hawalah, A., & Hussain, A. (2016). Comparing Oversampling Techniques to Handle the Class Imbalance Problem: A Customer Churn Prediction Case Study. IEEE Access, 4, 7940–7957. https://doi.org/10.1109/ACCESS.2016.2619719

Amirian, S., Gao, F., Littlefield, N., Hill, J. H., Jr., Plate, J. F., Pantanowitz, L., Rashidi, H., & Tafti, A. P. (2025). State-of-the-Art in Responsible, Explainable, and Fair AI for Medical Image Analysis. IEEE Access, 13, 58229–58263. https://doi.org/10.1109/ACCESS.2025.3555543

Arfaoui, N. (2026). Enhancing Machine Learning Algorithms for Imbalanced Data—A Case Study: Spam Detection. IEEE Access, 14. https://doi.org/10.1109/ACCESS.2026.3658624

Berengueres, J. (2024). How to Regulate Large Language Models for Responsible AI. IEEE Transactions on Technology and Society, 5(2), 191–197. https://doi.org/10.1109/TTS.2024.3403681

Biau, G. (2012). Analysis of a Random Forests Model. Journal of Machine Learning Research, 13, 1063–1095.

Breiman, L. (1999). Random Forests—Random Features. Technical Report 567, Department of Statistics, University of California, Berkeley.

Carballo-Mendívil, B., Arellano-González, A., Ríos-Vázquez, N. J., & Lizardi-Duarte, M. del P. (2025). Predicting Student Dropout from Day One: XGBoost-Based Early Warning System Using Pre-Enrollment Data. Applied Sciences, 15(16), 9202. https://doi.org/10.3390/app15169202

Cheng, K., Zhang, C., Yu, H., Yang, X., Zou, H., & Gao, S. (2019). Grouped SMOTE With Noise Filtering Mechanism for Classifying Imbalanced Data. IEEE Access, 7, 170668–170681. https://doi.org/10.1109/ACCESS.2019.2955086

Fu, Z., Ma, H., Wang, F., Dou, J., Zhang, B., & Fang, Z. (2024). An Integrated Framework of Positive-Unlabeled and Imbalanced Learning for Landslide Susceptibility Mapping. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 17. https://doi.org/10.1109/JSTARS.2024.3452182

Goswami, R., Garai, A., Sadhukhan, P., Ghosh, P., & Chakraborty, T. (2025). Shape Penalized Decision Forests for Imbalanced Data Classification. IEEE Access, 13. https://doi.org/10.1109/ACCESS.2025.3569523

Hu, J., & Szymczak, S. (2023). A Review on Longitudinal Data Analysis with Random Forest. Briefings in Bioinformatics, 24(2). https://doi.org/10.1093/bib/bbad002

Huang, Y. (2025). FS-SMOTE: An Improved SMOTE Method Based on Feature Space Scoring Mechanism for Solving Class-Imbalanced Problems. IEEE Access, 13, 148074–148082. https://doi.org/10.1109/ACCESS.2025.3597794

Jeong, J., Kahng, H., & Kim, S. B. (2026). Semi-Supervised Learning Under Extreme Class Imbalance in Wafer Bin Map Defect Pattern Classification. IEEE Access, 14. https://doi.org/10.1109/ACCESS.2026.3659740

Kim, D., Lee, W.-S., Ko, Y.-W., & Lee, J.-G. (2025). Separated and Independent Contrastive Semi-Supervised Learning for Imbalanced Datasets. IEEE Access, 13, 105712–105723. https://doi.org/10.1109/ACCESS.2025.3580738

Kim, K. (2021). Noise Avoidance SMOTE in Ensemble Learning for Imbalanced Data. IEEE Access, 9, 143250–143265. https://doi.org/10.1109/ACCESS.2021.3120738

Kim, S., Choi, E., Jun, Y.-K., & Lee, S. (2023). Student Dropout Prediction for University with High Precision and Recall. Applied Sciences, 13(10), 6275. https://doi.org/10.3390/app13106275

Kumar, R., Kim, Y.-W., & Byun, Y.-C. (2025). Hybrid Framework Combining Diffusion-Based Image Augmentation and Feature Level SMOTE for Addressing Extreme Class Imbalance. IEEE Access, 13. https://doi.org/10.1109/ACCESS.2025.3600622

Lee, C.-C., Comes, T., Finn, M., Pak, H., Hsu, C.-W., & Mostafavi, A. (2026). Roadmap Toward Responsible AI in Crisis Resilience and Management. IEEE Access, 14. https://doi.org/10.1109/ACCESS.2026.3651368

Matey-Sanz, M., Granell, C., & Mollineda, R. A. (2026). Responsible Integration of Generative AI in Software Engineering Education. IEEE Revista Iberoamericana de Tecnologias del Aprendizaje, 21.

Miftahushudur, T., Grieve, B., & Yin, H. (2024). Permuted KPCA and SMOTE to Guide GAN-Based Oversampling for Imbalanced HSI Classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 17, 489–505. https://doi.org/10.1109/JSTARS.2023.3326963

Mitchell, S., Potash, E., Barocas, S., D'Amour, A., & Lum, K. (2021). Algorithmic Fairness: Choices, Assumptions, and Definitions. Annual Review of Statistics and Its Application, 8, 141–163. https://doi.org/10.1146/annurev-statistics-042720-125902

Nazri, A., Agbolade, O., & Aziz, F. (2025). Hybrid Reinforcement Learning-Active Learning Framework for Real-Data Augmentation in Imbalanced Credit Scoring. IEEE Access, 13. https://doi.org/10.1109/ACCESS.2025.3608032

Niyogisubizo, J., Liao, L., Nziyumva, E., Murwanashyaka, E., & Nshimyumukiza, P. C. (2022). Predicting Student's Dropout in University Classes Using Two-Layer Ensemble Machine Learning Approach: A Novel Stacked Generalization. Computers and Education: Artificial Intelligence, 3, 100066. https://doi.org/10.1016/j.caeai.2022.100066

Pal, S. (2012). Mining Educational Data to Reduce Dropout Rates of Engineering Students. I.J. Information Engineering and Electronic Business, 4(2), 1–7. https://doi.org/10.5815/ijieeb.2012.02.01

Pessach, D., & Shmueli, E. (2022). A Review on Fairness in Machine Learning. ACM Computing Surveys, 55(3), Article 51. https://doi.org/10.1145/3494672

Putra, L. G. R., Prasetya, D. D., & Mayadi. (2025). Student Dropout Prediction Using Random Forest and XGBoost Method. INTENSIF: Jurnal Ilmiah Penelitian dan Penerapan Teknologi Sistem Informasi, 9(1), 147–157. https://doi.org/10.29407/intensif.v9i1.21191

Qin, H. (2025). Maximal Information Coefficient-Based Undersampling Method for Highly-Imbalanced Learning. IEEE Access, 13, 4126–4135. https://doi.org/10.1109/ACCESS.2025.3525475

Resende, P. A. A., & Drummond, A. C. (2018). A Survey of Random Forest Based Methods for Intrusion Detection Systems. ACM Computing Surveys, 51(3), Article 48. https://doi.org/10.1145/3178582

Ridwan, A., Priyatno, A. M., & Ningsih, L. (2024). Predict Students' Dropout and Academic Success with XGBoost. Journal of Education and Computer Applications, 1(2), 1–8.

Romsaiyud, W., Nurarak, P., Phiasai, T., Chadakaew, M., Chuenarom, N., Aksorn, P., & Thammakij, A. (2024). Predictive Modeling of Student Dropout Using Intuitionistic Fuzzy Sets and XGBoost in Open University. Proceedings of the 7th International Conference on Machine Learning and Machine Intelligence (MLMI 2024). https://doi.org/10.1145/3696271.3696288

Schenk, P. O., & Kern, C. (2024). Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production. AStA Wirtschafts- und Sozialstatistisches Archiv, 18, 131–184. https://doi.org/10.1007/s11943-024-00344-2

Shafiq, D. A., Marjani, M., Habeeb, R. A. A., & Asirvatham, D. (2022). Student Retention Using Educational Data Mining and Predictive Analytics: A Systematic Literature Review. IEEE Access, 10, 72480–72515. https://doi.org/10.1109/ACCESS.2022.3188767

Sharma, A., Singh, P. K., & Chandra, R. (2022). SMOTified-GAN for Class Imbalanced Pattern Classification Problems. IEEE Access, 10, 30655–30672. https://doi.org/10.1109/ACCESS.2022.3158977

Shakeel, K., & Butt, N. A. (2015). Educational Data Mining to Reduce Student Dropout Rate by Using Classification. Proceedings of the International Conference on Information and Communication Technologies.

Sowmya, P., & Vasudeva. (2026). Trust and Ethics in Chatbots: A Framework for Responsible AI Interactions. IEEE Access, 14. https://doi.org/10.1109/ACCESS.2026.3653763

Tsai, H.-C., Lee, M.-C., & Hsu, C.-H. (2025). Fault and Severity Diagnosis Using Deep Learning for Self-Organizing Networks With Imbalanced and Small Datasets. IEEE Access, 13. https://doi.org/10.1109/ACCESS.2025.3537659

Wang, J., & Awang, N. (2024). MKC-SMOTE: A Novel Synthetic Oversampling Method for Multi-Class Imbalanced Data Classification. IEEE Access, 12, 196929–196938. https://doi.org/10.1109/ACCESS.2024.3521120

Wang, Z., Xie, J., & Zhang, J. (2024). A Robust Enhanced Ensemble Learning Method for Breast Cancer Data Diagnosis on Imbalanced Data. IEEE Access, 12, 189776–189788. https://doi.org/10.1109/ACCESS.2024.3516376

Xing, Y. (2024). The Influence of Responsible Innovation on Ideological Education in Universities Under Generative Artificial Intelligence. IEEE Access, 12, 133008–133017. https://doi.org/10.1109/ACCESS.2024.3459469

Yan, Y., Liu, R., Ding, Z., Du, X., Chen, J., & Zhang, Y. (2019). A Parameter-Free Cleaning Method for SMOTE in Imbalanced Classification. IEEE Access, 7, 23537–23548. https://doi.org/10.1109/ACCESS.2019.2899467




DOI: http://dx.doi.org/10.35671/telematika.v19i2.3425

Refbacks

  • There are currently no refbacks.


 



Indexed by:

   

Telematika
ISSN: 2442-4528 (online) | ISSN: 1979-925X (print)
Published by : Universitas Amikom Purwokerto
Jl. Let. Jend. POL SUMARTO Watumas, Purwonegoro - Purwokerto, Indonesia


Creative Commons License This work is licensed under a Creative Commons Attribution 4.0 International License .