The study tested medium-sized, openly available large language models on a dataset of 142 post-doctoral fellowship applications submitted to the Swedish Medical Council in 1994. The researchers compared the models’ scores with the average scores given by expert reviewers and with the agreement among reviewers.

The models’ scores were moderately related to one another but only weakly related to the experts’ average scores. The authors say the systems might assist with application triage or help break ties, rather than replace expert review.