Health & Medicinearticle2026-08-08

Clinical liability cases reveal a coupling between legal defensibility and medical procedure escalation in large language models

Open access0 citations

Abstract

Large language models (LLMs) are increasingly used for clinical decision support, yet their resource utilization implications remain largely unstudied. We introduce CLAD (Clinical Legal Accountability Dataset), evaluating fifteen LLMs on 198 US medical malpractice cases to assess legal defensibility (alignment with court-determined standards of care) and associated medical procedure costs across 3072 simulated consultations. Here we show striking variation in defensibility: GPT-5.2 achieves a mean defensibility score of 0.71 (70% of consultations defensible, i.e., addressing the primary court-endorsed action) versus 0.34 for GPT-4o (32% defensible). Models achieving higher defensibility systematically recommend more medical procedures at substantially higher costs (Spearman ρ = 0.95): GPT-5.2 averages 9.3 procedures per consultation at $1,118 estimated Medicare cost, compared to 1.3 procedures and $221 for GPT-4o (5.1 × difference). This cost-defensibility coupling intensifies across model generations within the same provider family. Concise system prompting reduces response length by 59% but costs by only 5%, suggesting procedure escalation is embedded in model reasoning rather than verbosity. The coupling is not explained by training-data contamination (pre-2023 vs. post-2023 cases show no differential advantage for newer models), persists within case-severity strata, and replicates on an independent set of UK court cases (ρ = 0.94 across the same 15 models). These findings have substantial implications for healthcare costs as AI adoption scales, and the mechanisms driving this coupling, whether improved clinical reasoning, training-induced omission aversion, or a combination of both, warrant further investigation. Doctors are increasingly using artificial-intelligence chatbots to help make clinical decisions, and we wanted to know how relying on their recommendations would hold up to the legal standard of care expected of physicians, in other words, whether following the advice would help avoid medical malpractice. We studied this using real medical malpractice court cases, in which judges have already ruled on what appropriate care should have looked like. We asked fifteen artificial-intelligence systems to act as the doctor in each case and checked whether their advice met the court’s standard. We found that newer and more capable systems gave more legally defensible advice, but that they achieved this by ordering far more tests and referrals, at much higher cost. As hospitals adopt these tools, safer artificial-intelligence advice may therefore substantially increase health-care costs. Aran et al. evaluate 15 large language models on 198 US medical malpractice cases to assess legal defensibility (alignment with court-determined standards of care) and associated procedure costs for simulated medical consultations. Defensibility varies, with models that achieve higher defensibility recommending more procedures at higher costs.

// Source

View paper (DOI)Open access versionOpenAlexCommunications MedicinePublished 2026-08-08

Authors: Dvir Aran, Ronen Perry, Shahar Shelly

Institutions: University of Haifa, Rambam Health Care Campus, Technion – Israel Institute of Technology