AI & Computingarticle2026-08-11

Enhancing lung cancer subtype classification using histopathology foundation models in weakly-supervised and unsupervised learning frameworks

Open access0 citations

Abstract

The lack of demographically diverse datasets remains a major barrier to the deployment of computational pathology models in real-world clinical settings. To address this gap, we introduce and release the Indian Pathology Dataset for Lung Cancer (IPD-LUNG), a curated histopathology dataset comprising 311 high-resolution hematoxylin and eosin (H&E) whole slide images (WSIs) of Non-Small Cell Lung Cancer (NSCLC) from an Indian cohort. Using this dataset, we develop and evaluate a comprehensive computational framework for lung cancer subtyping based on histopathology foundation model representations. We investigate two complementary classification pipelines: a weakly supervised multiple instance learning (MIL) framework and an unsupervised phenotype discovery followed by supervised classification pipeline that constructs slide-level representations using Leiden clustering. On the IPD-LUNG dataset, the weakly supervised MIL framework achieves an accuracy of 0.916 and AUC of 0.912, while the phenotype discovery pipeline attains a competitive accuracy of 0.867 and AUC of 0.903, despite relying on label-free feature construction. Cross-dataset evaluation with TCGA-NSCLC further demonstrates strong generalization across populations and acquisition settings. Together, these results highlight the robustness of foundation model representations and demonstrate a scalable and interpretable framework for lung cancer subtyping.

// Source

View paper (DOI)Open access versionOpenAlexScientific ReportsPublished 2026-08-11

Authors: Karan V. Padariya, Piyush Singh, Shantveer G. Uppin, Monalisa Hui, C. V. Jawahar, P. K. Vinod

Institutions: International Institute of Information Technology, Hyderabad, Nizam's Institute of Medical Sciences