External Validation of Clinical Risk Scores and Machine Learning Models for Predicting 30-Day Cardiovascular Risk After Noncardiac Surgery: The PERICARE Study
Abstract
Background: Machine learning (ML) has emerged as a promising approach for preoperative cardiovascular risk prediction; however, the generalizability of ML models across institutions remains uncertain. Moreover, comprehensive head-to-head comparisons between ML algorithms and established clinical risk scores for predicting 30-day major adverse cardiac events (MACE) are scarce. We therefore evaluated the performance and external transportability of multiple ML models across independent centers and compared their predictive accuracy with validated benchmark clinical risk scores. Methods: In a site-separated two-center cohort (derivation n = 707, 27 MACE; external validation n = 378, 38 MACE), ten algorithms trained on preoperative variables were externally validated without refitting and benchmarked against the American university of Beirut-HAS2 (AUB-HAS2), American society of anesthesiologists (ASA), and revised cardiac risk index (RCRI). We assessed AUROC, calibration, Brier score, and decision-curve net benefit, with paired bootstrap comparisons, DeLong testing, IDI, and NRI. Results: Among the ML models, no single algorithm consistently outperformed the others across all performance metrics. Naive Bayes achieved the highest external discrimination (AUROC 0.738, 95% CI 0.668–0.804) but showed poor calibration and threshold-dependent clinical utility. Gradient Boosting showed the most favorable balance of discrimination and calibration slope (AUROC 0.707; calibration slope 0.991), although absolute risk remained underestimated in external validation, whereas HistGradient Boosting yielded the best overall probability prediction, with the lowest Brier score (0.087) and the greatest decision-curve net benefit at clinically relevant risk thresholds. However, in external validation, no ML model demonstrated statistically significant AUROC superiority over the AUB-HAS2 score (all p > 0.05), although several models modestly outperformed the RCRI. Conclusions: In this site-separated external validation, ML models showed metric-dependent performance but no discrimination advantage over the AUB-HAS2 index. Given low event counts, flexible-model results are hypothesis-generating. These findings provide a cautionary, reproducible benchmark; local recalibration and prospective evaluation are prerequisites before clinical deployment.
// Source
Authors: Aslan Erdoğan, Şeyma Yeşil, Gamze Gençol Akçay, Ufuk S Halil, İhsan Demirtaş, Ezgi Alp, Berat Erdem, Fatma Ekici, Tuğba Öztürk, İpek Göker, Mustafacan Kılıçoğlu, Hüseyin Akgün, Yusuf İnci, Utku ULUKÖKSAL, Furkan Fatih Yücedağ, Alihan Ayata, Baran Yılmaz, Taylan Akgün, Selim Topçu, Can Yücel Karabay
Institutions: University of Health Sciences Antigua, Sağlık Bilimleri Üniversitesi, University of Health Science, Medical Park Gaziantep Hospital, Sakarya Eğitim ve Araştırma Hastanesi, Yeditepe University, Gaziantep Children's Hospital