Researchers reviewed studies that tested artificial intelligence systems for finding fractures in human radiographs, commonly called X-rays, and computed tomography scans. They focused on studies that tested the systems on external data and compared them with clinicians. Nineteen studies met the criteria: 13 involving radiographs and six involving CT scans.

When complete data were available, AI showed pooled sensitivity of 86% and specificity of 84% for radiographs, and pooled sensitivity of 93% and specificity of 92% for CT scans. These results suggest that the systems often performed comparably to human readers, but the review rated the certainty of this comparison as very low and said the evidence is not yet sufficient for broad clinical implementation.