Bridging Clinical Practice and Research: Semi-Automated Extraction and Analysis of Parkinson’s Imaging Data from Real-World Setting
Abstract
Neuroimaging is essential for understanding Parkinson’s disease (PD) and other atypical parkinsonian syndromes. However, the scarcity of large-scale, accurately annotated datasets presents a critical challenge for training robust machine learning models and advancing translational research. To address this issue, we present a semi-automated extraction and NLP-driven labeling pipeline tested in a centralized multi-site clinical archive that extracts, curates, and validates neuroimaging cohorts directly from real-world clinical settings. Using this framework, we compiled a comprehensive brain imaging dataset of 2,944 subjects, including patients with PD, individuals with other extrapyramidal syndromes, and controls. To overcome the annotation challenge, we developed a semi-automated labeling strategy driven by natural language processing (NLP). This strategy extracts precise clinical labels from unstructured, specialist-authored radiological reports. We rigorously validated the integrity of the generated dataset through classical voxel-wise univariate analyses and region-based multivariate statistical learning methods. These analyses successfully identified significant structural alterations in regions previously associated with the condition, such as the thalamus, hippocampus, and cerebellum. Our findings suggest that routine clinical repositories can be systematically transformed into high-quality, research-grade datasets. This methodology provides a valuable resource for PD neuroimaging research and proposes a framework with the potential to scale for accelerating AI-driven biomarker discovery in clinical practice.
// Source
Authors: R. López, F. Segovia, F. J. Martinez-Murcia, J. Ramírez, T. Martin – Noguerol, F. Paulano-Godino, A. Luna, J.M. Gorriz
Institutions: Universidad de Granada, Instituto de Investigación Biosanitaria de Granada