AI & Computingpreprint2026-08-04

The Discriminating Power of Shape Comparison in Rank Space A Quantitative Calibration of the Ordering–Fitting Method

Open access0 citations

Abstract

Abstract Sort a set of quantities in descending order, fit the resulting (rank, value) curve with four families of functions — linear, logarithmic, exponential and power law — and take the one with the highest coefficient of determination as the "shape" of the series. This ordering–fitting practice recurs across biology, epidemiology, finance and scientometrics, yet it has almost never been calibrated: what, exactly, does an R² of 0.99 distinguish? This paper calibrates the method using closed-form correspondences together with Monte Carlo simulation. Most of what follows counts against the method. The main results. There is a closed-form one-to-one correspondence between shapes in rank space and four distributional families — uniform ↔ linear, exponential ↔ logarithmic, log-uniform ↔ exponential, Pareto(α) ↔ power law with rank exponent 1/α — verified term by term in the noiseless limit. Dynamic range yields a criterion that requires no fitting at all: for an exponential parent with sample size n, the base-10 span of the range is bounded above by roughly log₁₀(n·ln n), so an observed range exceeding that bound rules out the exponential parent outright; and the fitted decay rate in rank space obeys the identity k·n/ln10 = L, with relative error within ±2.2% across twelve combinations. Two independent calibrations of discriminating power land on the same order of magnitude: with eight points, R² ≥ 0.9785 carries a likelihood ratio of 2.17 between uniform and exponential parents, while the event "exponential shape wins" carries a likelihood ratio between 1.2 and 2.2 across parent families calibrated to the observed (n, range). The method is systematically biased toward the exponential shape: when the parent really is exponential, at n = 10 the wrong verdict (39%) is more than twice as frequent as the right one (18%). Head truncation — taking a "top 30" list — replaces the answer entirely, and the correspondence ceases to apply. The log-normal, which lies outside the correspondence, is classified as "exponential shape" roughly half the time and as "power law" the other half, and almost never as logarithmic. A six-item reporting specification follows from these results. Applied to a set of sixteen cross-domain series, none of the eight principal scenarios retains the status of a determination; all are demoted to records. Yet the exponential parent is still ruled out independently, by dynamic range alone, in four high-range scenarios — which is the strongest class of statement ordering–fitting can still deliver at these sample sizes.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-04

Authors: Qinfu Li