AI & Computingarticle2026-09-18

Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data

Open access1 citations

Abstract

Abstract Purpose Despite the importance of peer review for grant funding decisions, academics are often reluctant to conduct it. This can lead to long delays between submission and the final decision as well as the risk of substandard reviews from busy or non-specialist scholars. At least one funder now uses Large Language Models (LLMs) to reduce the reviewing burden but the accuracy of LLMs for scoring grant proposals needs to be assessed. Design/methodology/approach This article compares scores from a range of medium sized open-weight LLMs with peer review scores for a well-researched dataset, 142 Swedish Medical Council post-doctoral fellowship applications from 1994. Findings Whilst the LLM scores correlate moderately between each other (mean Spearman correlation: 0.34), they correlated weakly but positively and mostly statistically significantly with the average expert scores (mean Spearman correlation: 0.22). The highest rank correlation between expert scores and LLMs was 0.33 for Gemma 3 27 b based on proposal titles and summaries without their main texts, which is about half (56 %) of the correlation between reviewers. Research limitations The small sample size, old funding call and heterogeneous evaluation criteria all undermine the robustness of the analysis. Practical implications Despite the ability of LLMs to score grant proposals being quantitatively weaker than that of experts, at least in this special case, they may have role in application triage or tie-breaking. Originality/value This is the first assessment of the value of LLM scores for funding proposals.

// Source

View paper (DOI)Open access versionOpenAlexJournal of Data and Information SciencePublished 2026-09-18

Authors: Ulf Sandström, Mike Thelwall

Institutions: University of Sheffield, KTH Royal Institute of Technology