Core open science practices in major medical journals: development of an automatized tool based on a Large Language Model
Abstract
Background Open science practices are increasingly promoted to improve the transparency and reproducibility of biomedical research, and their large-scale monitoring has been identified as a priority by initiatives such as the UNESCO Recommendation on Open Science. Existing automated tools cover only narrow subsets of practices. We developed and validated an automated tool, based on locally run open-weights large language models (LLMs), to extract a core set of 13 open science practices from published medical research articles. Methods We built a validation database of 303 research articles published in 2020–2021 in 10 major general medical journals, stratified across randomized controlled trials (RCTs), meta-analyses, and other designs. Duplicate manual extraction with expert adjudication served as the reference standard. The pipeline combined document conversion (GROBID), BM25-based chunk retrieval, and classification by three Llama models (3-8B, 3-70B, 3.3-70B) run entirely locally, with and without supplementary materials. Diagnostic performance (sensitivity, specificity, predictive values, F1) was computed against the reference standard. Results Usable outputs were obtained for 301/303 articles. Performance was heterogeneous across criteria rather than across models: practices reported in standardized statements (conflicts of interest, funding, author contributions, data sharing, pre-registration, reporting guidelines) were extracted reliably (F1 ≥ 0.75 for at least one model), whereas rare or inconsistently reported practices (code sharing, statistical analysis plan, protocol sharing, ORCID, open access status, preprints) remained difficult regardless of model size. Study-type classification exceeded 0.90 on all metrics. Larger models improved several criteria at a 6.6–7.5× computational cost, and adding supplementary materials did not improve performance, mainly due to conversion failures. Conclusions A fully local, open-weights LLM pipeline can reliably monitor a substantial subset of open science practices. Document conversion and layout heterogeneity, rather than language understanding, constitute the main barrier to scaling automated open science monitoring.
// Source
Authors: Constant Vinatier, Guillaume Freyermuth, Margaux Millour, Gauthieu Le Bartz Lyan, Sarah Buet, Perrine Lunel, Nicolas Robillard, Florian Naudet, Mathieu Acher
Institutions: Inserm, Centre National de la Recherche Scientifique, Institut Universitaire de France, Université de Rennes, Université Rennes 2, Institut de Recherche en Informatique et Systèmes Aléatoires, Institut de Recherche en Santé, Environnement et Travail