Biologyarticle2026-08-11

When Clinicaland Metabolomics Data Work Together:A Comparative Framework for Multimodal Disease Classification acrossMachine and Deep Learning

Open access0 citations

Abstract

Abstract Disease classification using clinical and metabolomics data increasingly relies on multimodal integration, yet the complementary and comparative contributions of these modalities remain poorly understood. Most existing frameworks prioritize predictive performance without systematically examining how data modalities and modeling paradigms influence classification outcomes. Consequently, the relative value of individual modalities versus their integration, particularly in terms of model behavior, robustness, and interpretability, remains poorly characterized across machine learning (ML) and deep learning (DL) approaches. Here, we present a modality-aware comparative framework that enables direct, side-by-side evaluation of clinical-only, metabolomics-only, and combined modeling strategies within a unified pipeline. Unlike existing tools designed primarily for multiomics integration, this framework explicitly assesses when and how each modality contributes value across diverse classification scenarios. It supports the efficient use of available data while enabling a systematic comparison of model performance, stability, and interpretability across ML and DL methods. Rather than introducing a new classifier model, this work delivers a unified benchmarking workflow for evidence-based decision-making of modeling strategies in small-sample clinical metabolomics settings. Applied to two glomerulonephritis cohorts representing clinically-driven and metabolomics-driven classification settings, the framework revealed that the dominant contributing modality shifted by scenario, reflecting differences in disease context. These scenarios reflect realistic situations where discriminative signals arise from clinical variables, metabolic alterations, or their combination, as well as more complex cases involving subclass discrimination with overlapping profiles. While several models achieved comparable predictive accuracy, they differed in feature ranking stability, sampling sensitivity, and tendency to overfitting. Overall, this framework facilitates transparent, evidence-based selection of modeling strategies and data modalities suited to data complexity, sample size, and research objectives. Source code is available at https://github.com/kwanjeeraw/MMFramework.

// Source

View paper (DOI)Open access versionOpenAlexACS Measurement Science AuPublished 2026-08-11

Authors: Kwanjeera Wanichthanarak, Kassaporn Duangkumpha, Nichapa Kleebkomut, Pairash Saiviroonporn, Trongtum Tongdee, Yongyut Sirivatanauksorn, Chagriya Kitiyakara, Sakda Khoomrung

Institutions: Mahidol University, Kidney Disease Association of Thailand