Health & Medicinepreprint2026-08-18

The Verdict Is a Function of the Slice: Auditing a Published Quantization Superiority Claim

Open access0 citations

Abstract

A published comparison is read as deployment guidance, and what that reader needs is that its verdict holds on the files they are about to download. Two moves are supposed to make such a reading trustworthy: fix what will be read before the data exist, and reproduce it on a second inference engine. We price both, auditing whether one published sentence --- that AWQ "consistently outperforms" GPTQ "across different model scales (7B-70B)" --- transfers to released artifact pairs inside a deployment envelope declared in advance, one consumer-class 12 GB accelerator. Neither returns what it is bought for. Five of the eight downstream tasks we froze did not resolve the 1-point effect we had ourselves declared meaningful, the task we pre-registered among them (1.187--1.221 points of half-width); the second serving stack changed no verdict --- a null belonging to our aggregation convention as much as to the engines --- while the downstream criterion run on it alone accounts for 55.78% of the full audit's measured GPU seconds, and two criteria on one engine return the full audit's verdict on every claim cell at 40.14% of that cost. The sentence stands under a perplexity table over two generations of one model family. Carried on that same criterion and corpus to released 4-bit pairs of the very checkpoints that table ranges over, it returns its own direction at one of them on both stacks, falls to under a third of the frozen effect and survives correction on neither stack at a second, is indistinguishable from zero at a third, and gives opposite directions on the two stacks at a fourth --- 7B and 13B of a range quantified to 70B, and no statement about the measurement that table reports. What buys an answer is the second measurement regime --- criterion, evaluation data and unit of analysis moving together. Read as deployment guidance and crossed with engine --- task and aggregation convention being nested inside the downstream regime rather than factors beside it --- the sentence meets two regimes returning opposite conclusions on one released pair at a scale its quantifier covers, in a family that table does not contain, both clearing an effect fixed before any confirmatory data existed and both separating from zero under a permutation of the condition labels: on the audited papers' own criterion corpus, perplexity places AWQ behind GPTQ by 0.993% and 0.920% on two serving stacks matched at the kernel class, while LAMBADA --- one of the 3 tasks resolving the effect there --- places AWQ ahead by 3.136 and 3.272 accuracy points. What was fixed in advance is the margins, the estimator and the primary task; LAMBADA is not that task, and reading all eight was decided after the frozen single-task result, so the accuracy side of this reversal is a selection made on the outcome and the reversal is reported as an existence statement rather than as a pre-registered test. The direction turns over at two further layers of the instrument, in preliminary readings entering no verdict: on a second evaluation corpus --- disjoint from the one calibration recipe either lineage publishes, the others disclosing none --- all four perplexity arms change sign, and a tokenizer packaging setting outside either algorithm's specification moves two of them from excluding zero to containing it. On one of the lineage pairs a change of evaluation corpus moves the comparator's perplexity from 5.73 to 11,280 while leaving the claimed method between 24.14 and 25.00, on both stacks, every window scored. Whether such a sentence transfers therefore depends on which slice is read, on what that slice resolves, and --- across the checkpoints of its own lineage we measure --- on which of them a reader lands: one slice, including one fixed in advance, decides it in neither direction.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18

Authors: Guan-Yuan Chen, Ya-Fen Yeh

Institutions: National Tsing Hua University, North Carolina Exploring Cultural Heritage Online