Extraction-reasoning gaps in open-source LLM grading of liver pathology reports
Abstract
Large language models (LLMs) may reduce the manual effort required to convert narrative pathology reports into structured research data, but it remains unclear whether their performance reflects extraction of explicit diagnostic labels or genuine histopathological reasoning. We evaluated 28 open-source LLM configurations on 268 Chinese liver pathology reports from patients with metabolic dysfunction-associated steatotic liver disease (MASLD). Each configuration was assessed under two conditions: extraction from complete reports containing diagnostic conclusions and reasoning from microscopy descriptions after diagnostic conclusions were removed. Models graded fibrosis, inflammation, and steatosis across 10 repeated prompt-sampling runs. In the clinically realistic extraction condition, the strongest low-error models achieved high performance for fibrosis (Qwen3-14B: macro F1 0.928 ± 0.022; accuracy 0.961) and steatosis (GLM-4-32B: macro F1 0.961 ± 0.009; accuracy 0.956), whereas inflammation remained challenging. All tasks showed substantial extraction-reasoning gaps: fibrosis 0.327 (95% CI, 0.286–0.367), inflammation 0.221 (95% CI, 0.161–0.280), and steatosis 0.254 (95% CI, 0.230–0.279). Chinese-native model families showed an exploratory language advantage, especially at smaller scales. Reasoning-enhanced models often suffered from format failures. Open-source LLMs therefore appear promising for structured extraction from pathology reports, but reliable diagnostic reasoning from descriptive microscopy alone remains limited.
// Source
Authors: Chengying Xu, Hou Qiuchen, Lu Zhonghua
Institutions: Wuxi People's Hospital, Jiangnan University, Wuxi Fourth People's Hospital, Nantong University, Affiliated Hospital of Nantong University