Health & Medicinearticle2026-08-17

Accuracy assessment of Chinese large language models in psoriasis management: a multicenter expert evaluation study

Open access0 citations

Abstract

Psoriasis patients in China face severe barriers to disease knowledge acquisition and limited access to medical resources, creating an urgent demand for reliable, evidence-based patient education tools. This multicenter expert evaluation study aimed to systematically assess the quality and clinical applicability of mainstream Chinese large language models (LLMs) in answering psoriasis-related patient inquiries. A total of 40 clinically representative high-frequency questions covering 5 core domains (etiology, triggers, treatment, management, psychosocial impact) were curated from 365 Questions on Psoriasis (compiled by 109 Chinese psoriasis experts) by 9 board-certified dermatologists. Four mainstream Chinese LLMs (DeepSeek-R1, DeepSeek-V3, GLM-4, Qwen3-Plus) were evaluated via a double-blind expert scoring protocol. Responses were independently rated by 9 dermatologists using a 10-point Likert scale across 4 equal-weight dimensions: accuracy, completeness, clarity, and safety. The overall mean scores of the 4 LLMs ranged from 7.30 ± 0.47 to 8.84 ± 0.42, with single-question scores ranging from 4.88 to 9.20. Qwen3-Plus achieved the highest overall score (8.84 ± 0.42), significantly outperforming the other 3 models (all P < 0.001), followed by DeepSeek-V3 (8.36 ± 0.41), DeepSeek-R1 (8.19 ± 0.48), and GLM-4 (7.30 ± 0.47, the lowest, all P < 0.001 vs. other models). A statistically significant difference was also found between DeepSeek-R1 and DeepSeek-V3 (Bonferroni-adjusted P = 0.010). All models avoided life-threatening misleading content, and 87.5% of responses proactively emphasized the necessity of physician consultation. However, 12.5% of responses deviated from evidence-based guidelines, 80% of which were concentrated in complex biologics-related topics, including incorrect claims of psoriasis cure by biologics, inconsistent descriptions of marketed biologics in Chinese mainland, and non-standard recommendations for special populations (pregnancy/lactation). Chinese LLMs show preliminary potential as supplementary tools for standardized psoriasis patient education on basic disease topics, with generally safe information output and appropriate guidance to seek professional medical care in most responses. However, significant performance heterogeneity across models, non-negligible guideline deviations in specialized clinical topics (especially biologics-related content), and low absolute inter-rater agreement of the scoring system indicate that the clinical applicability and reliability of current LLMs in complex psoriasis management scenarios are still limited, and they cannot replace professional dermatological advice.

// Source

View paper (DOI)Open access versionOpenAlexScientific ReportsPublished 2026-08-17

Authors: C Y Chen, Hao Huang, Bo Yu, Xiaoping Hu, Jialiang Shi, Fangpei Wu, Keren Zhou, Yufang Yao, Xingyun Zhao, Jieyi Wang, Zhuoxuan Chen, Jingwen Wu, Jianjun Huang, Xiaoli Li

Institutions: University of Hong Kong, University of Hong Kong - Shenzhen Hospital, Shenzhen University, Hangzhou First People's Hospital, Westlake University, Hong Kong University of Science and Technology, Peking University Shenzhen Hospital, Changsha Central Hospital, Second Affiliated Hospital of Xi'an Jiaotong University, Aerospace Center Hospital