AI & Computingarticle2026-09-06

"Reproducibility package: Predictive Similarity Does Not Imply Selection Reproducibility"

Open access0 citations

Abstract

Feature selection and class-imbalance handling are routinely combined in medical tabular classification, yet judged by predictive performance alone. Whetherresampling changes the reproducibility of the selection itself is left unmeasured, although selected variables are often reported as clinical findings. Thisstudy treats imbalance handling as a controlled pipeline factor. Every data-level sampler is paired with a matched no-resampling condition on the same validation split under a training-only contract, and reproducibility is quantified by a chance-corrected estimator over fifty outer fits. A sequential rule treats predictive competitiveness as a hard constraint and ranks qualifying candidates by reproducibility. Four selector families and five imbalance conditions were evaluated on six binary medical datasets, with imbalance ratios from 1.6 to 14.6. Of sixty matched external contrasts, twenty-three produced material reproducibility changes, twenty of them reductions. In sixteen, a material loss coexisted with a paired predictive difference below the prespecified tolerance, a pattern invisible to a purely predictive evaluation. Only two of seven prespecified omnibus tests survived multiplicity correction. The rule changed the recommended pipeline in five of six datasets. Predictive similarity therefore does not imply comparable selection reproducibility, and measuring both turns an unrecognised cost of resampling into an explicit pipeline choice.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-09-06

Authors: Francisco Ndong Ndong Obono, Zhang Jue

Institutions: Yulin University