Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms
Abstract
Abstract Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January–December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss’ κ; inter-model agreement via Cohen’s κ and McNemar’s test; and each LLM’s majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss’ κ 0.837–0.860). Pairwise inter-model agreement was asymmetric: ChatGPT–Gemini behaved near-identically (Cohen’s κ = 0.850, 95% CI 0.71–0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity ( p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS ( p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.
// Source
Authors: Anish Narayan, Frederick Mariajoseph, Malik Farooq, Adrian Praeger, Ronil Chandra, Idrees Sher, Lee‐Anne Slater, Calvin Gan, Andrew Gauden, Hamed Asadi, Justin Moore
Institutions: Monash University, Monash Health, Austin Health