Reproduced Numbers, Unlicensed Claims: Preregistered Acceptance Contracts for Post-Training Compression
Abstract
Reproducing a published table does not settle whether the sentence it supports was ever specified tightly enough to be accepted or rejected as written. We study that gap through the acceptance contract---the six fields a quality-preservation claim must fix before it is testable (estimand, scope, comparator, tolerance provenance, variation distribution, decision rule)---and an executable instrument for it: preregistered manifests compiled into statistical gates whose multiplicity budget sits inside the decision rule, calibrated end to end on enumerated simulator arms (worst of six preregistered boundary arms 0.0526, simultaneous 95% bound 0.0569), every verdict replaying from a released de-identified per-item derivative. On a development panel---one verbatim claim from each of six of the most-cited post-training compression papers, inside a pre-declared deployment envelope---0/6 supply a complete contract, and no verdict names more than an operating point: AWQ's superiority is licensed under the auditor's contract at both bit widths, as is Wanda's magnitude-pruning clause (its whole sentence Inconclusive) and KIVI's "up to 2%"---the panel's one author-authorised, coding-contested tolerance---while SparseGPT and an auditor-derived GPTQ proposition are adverse at every preregistered tolerance yet concordant with the papers' printed losses. The census is author-coded with outcomes visible and no second human coder; unmeasured quantifiers abstain rather than carry a verdict. We then froze the procedure whole under a time-stamped commit and applied its front-end policy prospectively to a family the audit had never touched: of 58 low-rank/SVD candidates retrieved under three frozen queries, zero are eligible under its surface-evidence cap, mostly at a single field---a comparator that binds to no baseline arm---one driver being the frozen mapping's own resolution, reported rather than amended. The two censuses are one measurement at two scales: the claims of this literature are not yet written to be checked. We release the engine, manifests, calibration and coding artifacts, challenge files, the complete zero-eligible census and the replayable derivative---a prototype for finite, encodable componentwise claims, not certification infrastructure.
// Source
Authors: Guan-Yuan Chen, Ya-Fen Yeh
Institutions: National Tsing Hua University, North Carolina Exploring Cultural Heritage Online