Health & Medicinearticle2026-08-14

An XGBoost-based predictive framework for diabetes mellitus multi-classification

Open access0 citations

Abstract

Abstract Diabetes mellitus is a persistent metabolic condition that requires accurate and early diagnosis to prevent severe complications. This paper proposes an Extreme Gradient Boosting (XGBoost)-based predictive framework for multi-classification of diabetes mellitus into non-diabetic, pre-diabetic, and diabetic classes. After standardization and exclusion of non-clinical identifiers, Duplicate clinical records were removed from the original dataset, leaving 826 unique records. Comprehensive preprocessing pipeline used a stratified 70:30 train-test split and five-fold cross-validation; scaling and resampling were performed only within training partitions. Experimental results on the original dataset XGBoost achieved an accuracy of 99.60%. Both Random Over Sampling (ROS) and Syntenic minority over sampling technique (SMOTE) have also achieved 99.60% accuracy but provided improved generalization at the expense of higher computational cost. In contrast, Random Under Sampling (RUS) and Cluster Centroids (CC) reduced accuracy to 92.74% and 90.32%, respectively due to information loss. Across five folds, the original XGBoost model achieved 98.79 ± 1.13% accuracy. Benchmarking against Logistic Regression, Random Forest, Support Vector Machine, Decision Tree, and K-Nearest Neighbors showed that XGBoost provided the strongest performance. These findings highlight the effectiveness of XGBoost while emphasizing classification accuracy in multiclass diabetes prediction systems.

// Source

View paper (DOI)Open access versionOpenAlexScientific ReportsPublished 2026-08-14

Authors: Moataz M. El Sherbiny, Mohamed G. Abdelfattah, Ali E. Takieldeen, Hossam El-Din Mostafa, Hala B. Nafea