Health & Medicinearticle2026-08-07

ECG-based detection of occlusion myocardial infarction: a dedicated deep neural network versus multimodal large language models and physicians — a retrospective diagnostic accuracy study

Open access0 citations

Abstract

Abstract Background Accurate electrocardiogram (ECG) interpretation for acute coronary occlusion is a critical, time-sensitive task in emergency care. The field is increasingly shifting from the ST-segment elevation myocardial infarction paradigm to the broader concept of occlusion myocardial infarction (OMI). This study compared the diagnostic accuracy of a dedicated occlusion-detection deep neural network, Queen of Hearts (QoH), two general-purpose multimodal large language models (LLMs), and emergency physicians for ECG-based OMI detection. Response consistency and confidence calibration of the LLMs were also evaluated. Methods This retrospective diagnostic accuracy study used 36 twelve-lead ECGs from patients referred for emergent coronary angiography, including 24 with angiographically confirmed acute coronary occlusion and 12 without. Five emergency medicine specialists, five emergency medicine residents, and two LLMs, ChatGPT 5.2 and Gemini 3 Pro, interpreted all ECGs under a standardized moderate-risk acute coronary syndrome scenario. The same ECGs were submitted to QoH. Each LLM evaluated the set five times in separate sessions, whereas QoH was queried once because it provides deterministic output. In the primary head-to-head analysis, each interpreter or interpreter group contributed a single decision per ECG. Areas under the curve (AUCs) were compared using the DeLong test, binary metrics using the exact McNemar test, and reliability using Fleiss’ kappa. Results QoH showed the highest discrimination for OMI (AUC 0.96), followed by ChatGPT (0.81) and Gemini (0.62). QoH also achieved the highest sensitivity (95.8%, missing one of 24 occlusions) and accuracy (86.1%), with significantly higher sensitivity than emergency specialists ( p = 0.016). Among physicians, specialists performed best (accuracy 75.0%; sensitivity 66.7%; specificity 91.7%). ChatGPT showed high specificity (91.7%) but low sensitivity (54.2%), whereas Gemini performed least well overall (accuracy 55.6%). The LLMs produced variable interpretations across repeated queries (Fleiss’ kappa 0.24–0.49), and Gemini showed marked overconfidence (Brier score 0.40). Conclusions In this OMI-focused study, QoH showed the highest discrimination and sensitivity for acute coronary occlusion. General-purpose LLMs showed clinically important limitations and do not currently support autonomous OMI diagnosis. These findings support further evaluation of task-specific models for ECG-based occlusion detection and suggest, at most, an adjunctive, human-supervised role for current general-purpose LLMs.

// Source

View paper (DOI)Open access versionOpenAlexBMC Emergency MedicinePublished 2026-08-07

Authors: Emin Hüseyin Akar, Kâmil Kokulu, Ekrem Taha Sert

Institutions: Aksaray University