What a slice-level benchmark certifies without the pixels: a label-file audit of seven public benchmarks, and a protocol-frozen random-sample screen of how often the check is reported
Abstract
Background. Volumetric images are often labelled and scored one slice at a time. Shortcut learning, acquisition confounding and the inflation caused by slice-level splits are documented, and a pixel-free positional baseline is itself prior art; what is unmeasured is how much of a published number such a baseline accounts for under a correct split, and how often anyone checks. Purpose. To quantify how much published slice-level performance is reachable from a benchmark's own label file with no images, to test whether the evaluation unit changes which method wins, and to estimate how often this literature reports any such check. Materials and Methods. Pixel-blind null models requiring only four commonly published columns — subject identifier, slice index, label, train/test assignment — were fitted on training slices, and the resulting score vector was read at both the slice and the patient level with subject-clustered bootstrap intervals. Seven public benchmarks (eight label files) were audited. A rank-inversion analysis on 21 method configurations across five NYU fastMRI cohorts compared between-unit rank disagreement against within-unit resampling noise. A protocol-frozen random-sample screen of a frozen 9,979-record PubMed frame (seed 20260729, four independent screeners) estimated reporting prevalence. Proportions carry Wilson 95% CIs. Results. On RSNA 2019 Intracranial Haemorrhage (752,802 slices, 18,938 patients) one pixel-blind score vector read 0.737 (95% CI: 0.735, 0.740) per slice and 0.453 (95% CI: 0.445, 0.461) per patient, an interval wholly below chance; all six official labels behaved the same way (gaps 0.205–0.307). Against peer-reviewed comparators the median share of the published margin over chance reached without pixels was 0.469 (IQR: 0.437, 0.490; n = 9 benchmark-arms), and two benchmark-arms did not fire at all (LUNA16, fraction −0.002; PI-CAI, positional baseline exactly 0.500 AUROC). Of 250 sampled papers, 91 were included and readable while 44 (32.6%; 95% CI: 25.3%, 40.9%) could not be obtained; none of the 91 reported a measured zero-image baseline (0/91, 0.0%; 95% CI: 0.0%, 4.1%; bounding interval 0.0%–32.6%), and none appeared in any of 345 coded records over 300 papers. No rank inversion survived (0 of 447 pairs), and on three of five cohorts two disjoint halves of the same subjects ranked methods in opposite orders. Conclusion. The unit at which a volumetric benchmark is read decides what its number means; against peer-reviewed numbers a pixel-blind model reached about half the reported margin over chance, and in a protocol-frozen sample of the literature no paper that could be read reported this check. Figures showing NYU fastMRI image data are omitted from this deposit under the fastMRI Data Sharing Agreement, which permits their use in academic publications but not in a redistributable archive. Their captions and the results they illustrate are unaffected.
// Source
Authors: Sathvik Loke, Ethan Thomas Johnson, Neeraj Sai Venkat Movva, Aditya Raut
Institutions: Illinois Mathematics and Science Academy, Green Valley High School