Review suggests AI can help spot fractures on X-rays and CT scans
The evidence came from externally tested systems, but study differences and patient-selection concerns limit confidence in the results.
Moderate evidenceReviewInterpret with caution
Medical disclaimer: This article summarizes research findings and is for informational purposes only. It is not medical advice.
Editorial illustration — not from the study.
Researchers reviewed studies that tested artificial intelligence systems for finding fractures in human radiographs, commonly called X-rays, and computed tomography scans. They focused on studies that tested the systems on external data and compared them with clinicians. Nineteen studies met the criteria: 13 involving radiographs and six involving CT scans.
When complete data were available, AI showed pooled sensitivity of 86% and specificity of 84% for radiographs, and pooled sensitivity of 93% and specificity of 92% for CT scans. These results suggest that the systems often performed comparably to human readers, but the review rated the certainty of this comparison as very low and said the evidence is not yet sufficient for broad clinical implementation.
What the review examined
The researchers conducted a systematic review of diagnostic-accuracy studies found in PubMed/MEDLINE, Embase, and the Latin American and Caribbean Health Sciences Literature database. Eligible studies tested AI directly on radiographs or CT images from living human participants, used external validation data, and included a formal comparison with clinicians. The researchers assessed risk of bias and applicability using the QUADAS-2 tool, summarized sensitivity and specificity, and performed exploratory random-effects meta-analyses when complete two-by-two data were available.
What the review concluded
Nineteen studies were included: 13 on radiographs and six on CT scans, covering fractures of the face, arms and legs, pelvis and hip, ribs, and spine. For the four radiograph studies with complete two-by-two diagnostic data, pooled sensitivity was 0.86, meaning the systems identified about 86% of fractures in the analyzed data, and pooled specificity was 0.84, meaning they correctly classified about 84% of non-fracture cases. For the three CT studies with complete data, pooled sensitivity was 0.93 and specificity was 0.92. The researchers reported that AI often performed comparably to human readers, but certainty was low for the pooled estimates and very low for AI-assisted interpretation compared with unaided human readers. These are study-level diagnostic accuracy results and do not establish that AI improves patient outcomes.
Who this may apply to
The findings are most relevant to researchers and healthcare systems assessing fracture-detection software on radiographs or CT scans in human clinical settings. They do not show that every AI tool, imaging type, fracture location, or healthcare setting will have the same performance, and they do not establish how using AI affects patient care or outcomes.
The significance
Accurate fracture detection could be a useful capability for imaging-support software, particularly because the review examined externally validated systems and compared them with clinicians. However, this review does not show that AI should replace clinicians, that it improves diagnosis in routine practice, or that it improves patient outcomes. The reported performance varied across studies and imaging types, and prospective testing within real clinical workflows remains important before the findings can be applied broadly to healthcare settings.
Limitations & evidence assessment
The review included 19 studies, but only four radiograph studies and three CT studies had enough complete data for pooled analysis. Study methods varied, and patient selection was the main concern: risk of bias was unclear in 11 studies and high in six. The review also reported low certainty for the pooled accuracy estimates and very low certainty for AI-assisted interpretation compared with unaided readers. The evidence came from externally validated studies, but prospective testing in real clinical workflows is still needed, and the abstract does not provide enough detail to determine how performance varied across all patient groups and settings.
Why this evidence level: This was a systematic review and meta-analysis of human diagnostic-accuracy studies, which generally provides more informative evidence than a single study. However, the review rated certainty as low for the pooled accuracy estimates and very low for comparisons between AI-assisted and unaided human readers, with substantial study limitations.
Evidence levels are editorial estimates derived from study metadata — they are not clinical appraisals.
// Source
Indian Journal of Musculoskeletal Radiology · 2026 · DOI: 10.25259/ijmsr_30_2026
A systematic review examined whether Passiflora edulis, commonly known as passion fruit, is linked with lower blood pressure. It included six preclinical studies and two human trials; the animal and laboratory findings were generally more extensive than the human evidence. The authors reported possible blood-pressure reductions but emphasized that the human studies were small, varied in the extracts used, and short in duration.
Researchers studied 96 people with high-risk, locally advanced rectal cancer in a single-center phase II clinical trial. The treatment combined chemotherapy, short-course radiation, and a targeted drug selected for each study cohort. Overall, 40% had a complete response, but the study did not establish how this approach compares with other treatments.
A prospective observational study of 143 people undergoing procedures for urinary stones compared pre-surgery urine cultures with cultures taken from the stones during surgery. Stone cultures were positive more often and were associated with postoperative sepsis, while pre-surgery urine culture results were not significantly associated with sepsis. Because this was an observational study, the findings show associations rather than cause and effect.