AI & Computingarticle2026-08-17

Comprehensive benchmarking of machine learning methods for risk prediction modelling from large-scale survival data: a UK Biobank study

Open access0 citations

Abstract

Predictive modelling is vital to guide the prevention of incident disease (e.g., risk-based initiation of preventive pharmacotherapies). Whilst large-scale prospective cohort studies and a diverse toolkit of available machine learning (ML) algorithms have facilitated such survival task efforts, choosing the best-performing algorithm remains challenging. Benchmarking studies to date focus on relatively small-scale datasets and it is unclear how well such findings translate to large datasets that combine omics and clinical features. We sought to benchmark eight distinct survival task implementations, ranging from linear to deep learning (DL) models, within the large-scale prospective cohort study UK Biobank (UKB). We compared discrimination and computational requirements across heterogeneous predictor matrices and endpoints. Finally, we assessed how well different architectures scale with sample sizes ranging from n = 5,000 to n = 250,000 individuals. Our results show that discriminative performance across a multitude of metrics is dependent on endpoint frequency and predictor matrix properties, with very robust performance of penalised Cox proportional hazards (Cox-PH) models (Ridge and Elastic Net both ranked first in 7 out of 15 instances when using the complete feature set). Of note, there are certain scenarios which favour more complex frameworks, specifically if working with larger numbers of observations and relatively simple predictor matrices; e.g. DL model performed best for cardiovascular risk prediction using clinical features (mean Harrell’s C 0.721 [95% confidence interval [CI] 0.719–0.722]). The observed computational requirements were vastly different, and we provide solutions in cases where current implementations were impractical. In conclusion, this work delineates how optimal model choice is dependent on a variety of factors, including sample size, endpoint frequency and predictor matrix properties, thus constituting an informative resource for researchers working on similar datasets. Furthermore, we showcase how linear models still display a highly effective and scalable platform to perform risk modelling at scale and suggest that these are reported alongside non-linear ML models.

// Source

View paper (DOI)Open access versionOpenAlexJournal Of Big DataPublished 2026-08-17

Authors: Rafael R. Oexner, R Schmitt, Hyunchan Ahn, S. S. Khawaja, R. A. Shah, A. Zoccarato, K. Theofilatos, Ajay M. Shah

Institutions: University College London, British Heart Foundation