Five published signals for updating meta-analyses, applied and evaluated: an exploratory study of when they can be used at all
Abstract
Background. Five statistical methods for deciding whether a systematic review has gone out of date were published between 1999 and 2007. They have been run side by side exactly once, on 80 Cochrane reviews, where two flagged nothing and the three that discriminated agreed at Kappa = 0.14. No reusable software implementation of any of them appears to exist, so the comparison has not been repeated. Methods. We implemented all five in an open-source R package and applied them to historical evidence in two arms. Arm A sweeps the 17 meta-analyses in the metadat collection that carry per-study data and publication years, running all five detectors against three operational definitions of the pooled estimate having moved. Arm B assembles 6,686 consecutive version pairs from the 4,132 Cochrane reviews with more than one version, harvested from Europe PMC, and attaches the editorial outcome recorded in each version's authors' conclusions, supporting the three detectors that need only pooled estimates. Results. In Arm A, 168 of 185 cuts (91%) had an already-significant prior meta-analysis, so barrowman and simulation — both of which require a non-significant prior — could be asked in only 4 and 5 of the 17 reviews. This was invisible to the published comparison because its cohort was selected for non-significance. The Ottawa method's effect criterion is a ratio of relative risk reductions whose denominator approaches zero as the prior effect approaches the null, so it fires on 64% of samples containing no change at all; its specificity falls to 0.14 on a null review where the other four hold at 1.00. The stability half of the sufficiency method is defined as the slope of a least-squares fit to the cumulative effect series; that slope has no valid null distribution and fired on 209 of 300 samples of unchanging evidence. In Arm B, 560 pairs carry both a comparable pooled estimate and authors' conclusions at each end; an automated screen separates likely conclusion changes 11.3-fold across strata. Under a protocol frozen before the evidence was opened, the three detectors evaluable from pooled estimates alone fire on 1.5%, 6.2% and 4.8% of the pairs they answer, against an event prevalence of 23% to 52% depending on the labelling; none approaches the pre-specified sensitivity of 0.60 under either. In 45% of the pairs whose conclusions changed, the pooled effect did not move at all, which caps any estimate-based detector at a sensitivity of 0.40 to 0.55 on this corpus. Conclusions. The five divide into two kinds of problem, and the distinction matters. Two are simply not defined once the prior meta-analysis is significant, which is the case at 91% of cuts here: that is a domain restriction rather than a defect, but it means they usually cannot be asked at all, and the one published comparison could not see this because its cohort was selected for the single condition under which they can. The other two failures are defects: one criterion is unstable by construction on exactly the reviews its method targets, and one statistic has no valid null distribution. The Arm A findings are exploratory, on 17 reviews with no held-out set, scored against operational targets that observe the pooled estimate moving rather than what any review team did; the Arm B evaluation is confirmatory and pre-specified, and its outcome is an automated label with measured error, not adjudicated truth. Beyond the individual failures, the corpus bounds the approach itself: nearly half of the editorial changes leave no trace in the pooled estimate, so the ceiling for this family of detectors sits below the bar that was set for it before any of them was run. The software, the sweeps and the corpus are reproducible from public sources.
// Source
Authors: Javier Núñez