Why grounded large language models fail without domain-specialized retrieval: an experimental scientometric study in solar physics
Abstract
Abstract Recent applications of large language models (LLMs) in scientometric analysis often assume that the underlying evidence space is neutral and pre-given. We challenge this assumption empirically by demonstrating that retrieval infrastructure constitutes a methodologically consequential component that systematically shapes model-based analytical outputs. Our framework combines a thematic core regime for semantic organization and retriever construction with a temporal holdout regime for frontier validation. It evaluates a generic SciBERT baseline, a domain-specialized dense retriever, and a lexical BM25 control over the same multimodal application corpus. We trained SciBERT-SolarPhysics-Search on 22,047 domain-adaptive pretraining documents and 43,314 supervised contrastive pairs, reaching a perplexity of 3.27 in the core training regime. The findings show that domain-specialized retrieval improves the quality of scientific retrieval, increasing MRR by +8.2%, Recall@10 by +2.9%, and Nearest-Centroid Accuracy by +51.0% over the generic dense baseline. Bootstrap confidence intervals support positive deltas for six of eight benchmark metrics, while NMI and Recall@100 remain directionally positive but not conclusive at the 95% level. Critically, these gains are most interpretable and defensible when we treat corpus roles, temporal separation, baseline comparison, and evidence traceability as explicit methodological controls rather than implementation details. These results suggest that reliable LLM-assisted scientometric analysis requires, at least in the solar physics domain studied in this work, explicit control over retrieval infrastructure, temporal regime, and evidence traceability. The findings provide a methodological template that other studies may adapt and test in additional scientific domains. Furthermore, treating information retrieval as a neutral pre-processing step risks systematically distorting the evidence space from which we extract analytical claims.
// Source
Authors: André Insardi, Andre Gradvohl
Institutions: Universidade Estadual de Campinas (UNICAMP), Escola Superior de Propaganda e Marketing