The beginning of the end for manual chart review: LLM-mediated database construction
Abstract
Manual clinical data abstraction is the reference standard for research databases but is labor-intensive, costly, and susceptible to human error. We evaluated the accuracy of a locally deployed open-source large language model (LLM) framework for automated extraction of structured kidney cancer data from unstructured clinical documentation. In this retrospective study, 8366 patients undergoing nephrectomy for suspected kidney or upper tract urothelial cancer between 2009 and 2024 were identified from a large academic health system. The LLM processed 130,509 clinical notes to generate 136,425 data elements across 14 operative, pathology, and radiology variables. Overall agreement with a manually curated reference database was 97.5%, with pathologic variables exceeding 98% agreement and Cohen’s κ values > 0.90. Manual arbitration of disagreements (2.5%) favored the LLM for 10 of 14 variables. These findings demonstrate that locally deployed open-source LLMs can accurately and efficiently generate large, structured clinical databases while preserving institutional data governance.
// Source
Authors: Jacob M. Knorr, Sahil Patel, Haya T. Abusafieh, Daniel Jevnikar, Rikhil Seshadri, Rishi Jonnalagadda, Jessica Yoon, Salim Younis, Nicolas Soputro, Gabriela Diaz, Betty Wang, Gagan Fervaha, Michal Ozery-Flato, Michal Rosen-Zvi, Rebecca Campbell, Steven Campbell, Venkatesh Krishnamurthi, Jihad Kaouk, Robert Abouassaly, Christopher J. Weight, Nicholas Heller
Institutions: Case Western Reserve University, West Virginia University, Cleveland Clinic, Cleveland Clinic Lerner College of Medicine, IBM Research - Haifa