Health & Medicinearticle2026-08-03

Assessing the clinical applicability of large language models for microwave ablation guidance

Open access0 citations

Abstract

Purpose To evaluate the performance of five language models (LLMs)—ChatGPT, Gemini, Meta, Claude, and DeepSeek—in providing clinically relevant advice on microwave ablation (MWA), a minimally invasive therapy for tumor treatment. Methods A total of 27 standardized questions covering seven thematic domains of MWA were presented to each LLM using a single-input approach. Responses were independently evaluated by three senior radiologists using three customized 4-point rating scales: Overall Quality Score (OQS), Understandability Score (US), and Implementability Score (IS). The mean of these metrics was calculated as the Mean Quality Score (MQS). Interrater reliability was assessed using intraclass correlation coefficients (ICC). Results 540 ratings were performed. All LLMs successfully generated relevant responses. ChatGPT achieved the highest MQS (3.65 ± 0.54), outperforming Claude (3.47 ± 0.59) and DeepSeek (3.47 ± 0.60), while Gemini (3.00 ± 0.55) and Meta (2.95 ± 0.61) scored significantly lower (p < 0.05). ChatGPT also ranked highest across all subcategories, particularly in Effectiveness and Outcomes (3.73 ± 0.44), Risks and Complications (3.89 ± 0.32), and Post-Procedure Care (3.89 ± 0.32). Interrater reliability was moderate to good (ICC = 0.56–0.75), confirming consistent evaluation. Conclusion All examined LLMs demonstrated the ability to provide meaningful advice on MWA. However, ChatGPT showed the best overall accuracy, clarity, and clinical applicability. Claude and DeepSeek achieved comparable results in specific areas. While these findings highlight the promising potential of LLMs in supporting clinical decision-making, human oversight and domain-specific validation remain essential.

// Source

View paper (DOI)Open access versionOpenAlexEuropean Journal of Radiology OpenPublished 2026-08-03

Authors: René Michael Mathy, Fabian Schmitz, Hans‐Ulrich Kauczor, Hyungseok Jang, Sam Sedaghat

Institutions: Heidelberg University, University Hospital Heidelberg, University of California, Davis