Health & Medicinearticle2026-09-04

External Validation of Clinical Risk Scores and Machine Learning Models for Predicting 30-Day Cardiovascular Risk After Noncardiac Surgery: The PERICARE Study

Open access0 citations

Abstract

Background: Machine learning (ML) has emerged as a promising approach for preoperative cardiovascular risk prediction; however, the generalizability of ML models across institutions remains uncertain. Moreover, comprehensive head-to-head comparisons between ML algorithms and established clinical risk scores for predicting 30-day major adverse cardiac events (MACE) are scarce. We therefore evaluated the performance and external transportability of multiple ML models across independent centers and compared their predictive accuracy with validated benchmark clinical risk scores. Methods: In a site-separated two-center cohort (derivation n = 707, 27 MACE; external validation n = 378, 38 MACE), ten algorithms trained on preoperative variables were externally validated without refitting and benchmarked against the American university of Beirut-HAS2 (AUB-HAS2), American society of anesthesiologists (ASA), and revised cardiac risk index (RCRI). We assessed AUROC, calibration, Brier score, and decision-curve net benefit, with paired bootstrap comparisons, DeLong testing, IDI, and NRI. Results: Among the ML models, no single algorithm consistently outperformed the others across all performance metrics. Naive Bayes achieved the highest external discrimination (AUROC 0.738, 95% CI 0.668–0.804) but showed poor calibration and threshold-dependent clinical utility. Gradient Boosting showed the most favorable balance of discrimination and calibration slope (AUROC 0.707; calibration slope 0.991), although absolute risk remained underestimated in external validation, whereas HistGradient Boosting yielded the best overall probability prediction, with the lowest Brier score (0.087) and the greatest decision-curve net benefit at clinically relevant risk thresholds. However, in external validation, no ML model demonstrated statistically significant AUROC superiority over the AUB-HAS2 score (all p > 0.05), although several models modestly outperformed the RCRI. Conclusions: In this site-separated external validation, ML models showed metric-dependent performance but no discrimination advantage over the AUB-HAS2 index. Given low event counts, flexible-model results are hypothesis-generating. These findings provide a cautionary, reproducible benchmark; local recalibration and prospective evaluation are prerequisites before clinical deployment.

// Source

View paper (DOI)Open access versionOpenAlexJournal of Clinical MedicinePublished 2026-09-04

Authors: Aslan Erdoğan, Şeyma Yeşil, Gamze Gençol Akçay, Ufuk S Halil, İhsan Demirtaş, Ezgi Alp, Berat Erdem, Fatma Ekici, Tuğba Öztürk, İpek Göker, Mustafacan Kılıçoğlu, Hüseyin Akgün, Yusuf İnci, Utku ULUKÖKSAL, Furkan Fatih Yücedağ, Alihan Ayata, Baran Yılmaz, Taylan Akgün, Selim Topçu, Can Yücel Karabay

Institutions: University of Health Sciences Antigua, Sağlık Bilimleri Üniversitesi, University of Health Science, Medical Park Gaziantep Hospital, Sakarya Eğitim ve Araştırma Hastanesi, Yeditepe University, Gaziantep Children's Hospital