Recovering Honest Inference: A Sparsity- and Skew-Aware Recalibration Doctrine for Pervasively Miscalibrated Small-Sample Chi-Square- and F-Referenced Tests
Abstract
Abstract: A practitioner who runs a test of independence, a goodness-of-fit test, a likelihood-ratio test, or an analysis of variance reads a single p-value against a chi-square or F reference as calibrated. In large samples it is, because each is asymptotically chi-square. But when expected counts are small, or groups are tiny and skewed, the exact small-sample law is a generalized chi-square the design has deformed, the asymptotic chi-square is only its limit, and the true level drifts off nominal, conservatively (wasting power) or liberally (creating false positives). The p-value cannot show which; that invisibility is the reporting problem. This paper states one recalibration move (reference the finite-sample law, or the variance-stabilized scale on which the statistic is near-normal, and report the routing descriptor) and subjects it to a five-step validation gate. Three procedures carried through the gate: in skewed ANOVA a variance-stabilized-location test restores calibration and, where both are valid, beats trimmed-means ANOVA on size-adjusted power; in goodness of fit Pearson's statistic is trustworthy far into sparsity while the likelihood-ratio statistic is the offender; and in small-sample regression the score test is the calibrated default and Wald the one to avoid. None is a new test. On public data, across 938 small-sample instances from 189 datasets, the two references disagree materially in about half and flip the 0.05 verdict in one in fourteen, directionally systematically. In the kidney data the raw-scale test misses a sex difference that the skew-aware test recovers (p = 0.092 versus 0.0085). Code is licensed MIT; documents CC BY 4.0.
// Source
Authors: William Dwyer
Institutions: University of Massachusetts Lowell