Development and benchmark validation of PubChat for PubMed-grounded multilingual biomedical literature retrieval
Abstract
The rapid growth of biomedical literature has rendered traditional systematic reviews unsustainable. Although large language models (LLMs) offer automation potential, citation fabrication and unreliable evidence discrimination remain critical barriers. Here, using PubChat as a PubMed E-utilities-grounded retrieval framework, we tested whether source-verifiable AI-assisted retrieval could preserve recall, criterion-driven relevance stratification, and multilingual accessibility in systematic-review benchmarking. The system comprises Phase I, hierarchical decomposition of the research question into five relevance levels, and Phase II, multi-round retrieval with embedding-based pre-filtering and three-round LLM verification. Validated against 20 Cochrane systematic reviews (585 ground-truth articles) across eight languages, PubChat was benchmarked against four general LLMs (GPT-5.2-Thinking, Gemini 3.0 Pro, Grok-4.1-Thinking, Qwen3-Max), one search-augmented retrieval tool (Perplexity-Sonar), and three specialized retrieval tools (Elicit, ASTA, and Consensus). PubChat produced no fabricated citations in this benchmark and showed distinct recall–precision profiles across its three operating modes. Among the evaluated configurations, PubChat-Broad achieved the highest observed recall and nDCG, whereas PubChat-Core achieved the highest observed F1- and F2-scores. In a user evaluation of 279 biomedical researchers across 18 countries, PubChat scored above 80/100 for reliability, innovation, efficiency, user experience, and self-reported preference over the evaluated alternatives (78% vs. specialized tools and 72% vs. manual search). An exploratory MIDE case study further illustrates how PubChat-derived corpora can be organized into evidence-traceable research-gap candidates for hypothesis prioritization. PubChat provides a benchmark-validated PubMed-grounded framework for source-faithful biomedical literature retrieval, relevance stratification, and structured evidence organization.
// Source
Authors: Xuenan Zhuang, Dan Cao, Ruoyu Chen, Yuxuan Wu, Qianjin Feng, Philip Yawen Guo, Gianvito Urgese, Shenghao Fu, Kun-Yu Lin, 林 芷君, Zedan Zhang, Ze Yuan, Enrico Borrelli, Feng Wen, Jiaxin Pu, Mengjie Li, Anyi Liang, Zhicong Xu, Zikang Xu, Danling Huang, T. Li, Ingrid Manguera Hojas, Mi Gui, Jiayi Huang, Jing Li, Xinyu Wang, Qi Chen, Yinzhou Zhong, Peibo Chen, Liang Zhang, Jinyun Jiang, Huanzi Lu
Institutions: University of Hong Kong, Peking University, Sun Yat-sen University, University of Alabama at Birmingham, Hong Kong Polytechnic University, Chinese University of Hong Kong, Southern Medical University, Wuhan University, Renmin Hospital of Wuhan University, Department of Medical Sciences, University of Turin, Guangdong Academy of Medical Sciences, Stomatology Hospital, Guangdong Provincial People's Hospital, CTO Hospital, Politecnico di Torino, Peking University First Hospital, Department of Public Health