Machine learning survival models for colorectal cancer: advantages of random survival forests over classification algorithms
Abstract
Colorectal cancer remains a major cause of cancer-related mortality worldwide, motivating the development of accurate prognostic models. Despite advances in machine learning for survival prediction, a common practice in the literature reformulates time-to-event outcomes as binary variables and applies standard classification algorithms, disregarding censoring and the temporal nature of survival data. In this study, we critically evaluate this practice and argue that it is unnecessary and methodologically concerning, given the availability of machine learning methods specifically designed for survival analysis that properly account for censoring while offering comparable interpretability. Using population-based data from colorectal cancer patients in the state of São Paulo - Brazil , we revisit a recent classification-based study that employed a classification XGBoost model and refit the same data using Random Survival Forests (RSF) under multiple splitting rules. Predictive performance was evaluated using Harrell’s concordance index and the Integrated Brier Score, and interpretability was assessed using SHAP values derived from a scalar risk quantity. RSF achieved stable performance and identified clinically meaningful predictors of mortality, while classification models systematically underestimated survival probabilities, particularly as censoring increased, a pattern confirmed through simulation studies. These findings reinforce the importance of using survival-specific machine learning methods for censored time-to-event data.
// Source
Authors: Rafael Herzog, Vinicius F. Calsavara, Agatha S. Rodrigues
Institutions: Universidade de São Paulo, Cedars-Sinai Medical Center, Universidade Federal da Paraíba, Universidade Federal do Espírito Santo