Health & Medicinearticle2026-08-28

Moderate discrimination but attenuated predictive precision: a multi-model external validation of prolonged postoperative length of stay models after lumbar interbody fusion in an independent cohort

Open access0 citations

Abstract

Prolonged postoperative length of stay (PLOS) following lumbar interbody fusion is associated with increased morbidity such as nosocomial infection and venous thromboembolism, healthcare costs, and impaired recovery. Several prediction models have been proposed to identify patients at risk for PLOS; however, none have undergone rigorous head-to-head external validation, and their generalisability across institutions remains uncertain. To externally validate and compare the performance of four previously published prediction models for PLOS after lumbar interbody fusion in an independent cohort. This retrospective observational study included consecutive adult patients undergoing primary open transforaminal lumbar interbody fusion at a high-volume tertiary spine center in northern Guangdong, China (2021–2024). Four published PLOS prediction models were externally validated using their original predictor sets without modification. Model performance was assessed across multiple domains, including discrimination (AUC), precision–recall characteristics, calibration, distribution of predicted risk, and clinical utility using decision curve analysis (DCA). PLOS was defined using the 75th-percentile approach consistent with the original studies. A total of 1,001 patients were included in the external validation cohort. Across the four models, discriminative performance was moderate, with optimism-corrected AUCs ranging from 0.63 to 0.72. Precision–recall analysis consistently demonstrated attenuated predictive precision, with AUC-PR values between 0.35 and 0.50 in cohorts with event prevalences of approximately 21–24%, indicating substantial false-positive rates at clinically relevant sensitivity levels. Calibration performance varied across models: while overall population-level calibration was acceptable, several models exhibited risk compression or instability in higher predicted-risk ranges, limiting reliable individual risk estimation. DCA demonstrated positive net benefit over treat-all strategies only within model-specific threshold ranges. However, the magnitude of benefit was generally modest, highly threshold-dependent, and insufficient to support stand-alone clinical decision-making. While the evaluated models demonstrated acceptable discrimination, their attenuated predictive precision, threshold instability, and sensitivity to population differences limit their standalone clinical applicability. Although several models demonstrated modest net benefit within selected threshold ranges, none provided sufficiently robust and threshold-invariant performance for stand-alone discharge planning or resource allocation decisions without prior local recalibration and contextual adaptation. Future efforts should focus on model updating, multicenter validation, and development of simplified clinical tools, to enhance real-world usability and integration into perioperative care pathways.

// Source

View paper (DOI)Open access versionOpenAlexEuropean Spine JournalPublished 2026-08-28

Authors: Xiangheng Dai, Z Chen, Chongle Huang, Yikai Xu, Jiali Zheng, Xinghong Zeng, Zijing Wu, Jiazhe Zhou, Linli Yan, Zhengyun Zhang, B Zhou, Qiang Wu

Institutions: Guangdong Medical College, Shaoguan University, Shaoguan Railway Hospital