Impact of prompt engineering and workflow modifications on the performance of a Delphi-based automated abstract screening system in a psychiatric systematic review
Abstract
Large language models (LLMs) can facilitate the labor-intensive screening of literature for systematic reviews and meta-analyses. Ensembling multiple LLMs to reach consensus on screening may allow locally deployed LLMs to achieve acceptable performance while reducing hallucination errors and maintaining privacy and reproducibility. The Delphi technique – in which panelists independently and anonymously explain their decisions, and then iteratively revise their judgements in response to aggregated responses from other panelists until reaching consensus – may provide a useful framework to formalize this process. To our knowledge, there are no published empirical evaluations on how changes in prompting impact the performance of automated screening when using multi-agent LLM ensembles in the medical domain. We tested the impact of different prompts and workflow modifications on the performance of an automated abstract screening workflow that integrated five LLMs into a Delphi ensemble for complex classification tasks. Our test dataset was a published systematic review of biomarkers of non-syndromic autism spectrum disorder, in which human screeners deemed 1,655 abstracts potentially eligible (“target abstracts”) among 4,745 identified by the search. We performed 12 types of experiments: five variations in the wording of the eligibility criteria section of the LLM prompt using moderate (“primary”), simplified, detailed, structured, or verbatim prompts, and seven modifications to the workflow’s technical features (i.e., the temperature, embedding model and adapter, few-shot example sampler, safety net, and ensemble settings). Our primary evaluation metric was the workflow’s recall (i.e., proportion of target abstracts detected by the workflow); secondary metrics included the work saved over sampling at 95% recall (WSS@95%, the effort saved by a screening system to identify 95% of the target abstracts compared to a random sampling strategy). Depending on the prompt and the workflow parameters, the screening system identified 59–99% of the 1,655 target abstracts. The WSS@95% ranged from 13 to 42%, indicating that changes in prompts and parameters moderately affect performance and efficiency. Simplified version of the prompt optimally balanced recall (98%) and WSS@95% (41–42%). Our study demonstrates the impact of prompt engineering and workflow modifications on an automated abstract screening workflow using Delphi-ensembled deployment of five open-weight LLMs. Future studies need to extend this work to automated data extraction to further advance the field of automated systematic review and meta-analysis.
// Source
Authors: Mirkamal Tolend, Ramzi Halabi, Kousai Ghaouari, Yvonne C.Y. Lau, Martin Alda, Arend Hintze, Benoit H. Mulsant, Abigail Ortiz
Institutions: Dalhousie University, Royal Ottawa Mental Health Centre, Southwestern Medical Center, The University of Texas Southwestern Medical Center, Centre for Addiction and Mental Health, National Institute of Mental Health, Dalarna University