Measure the Instrument First: Preregistered, Prediction-Sealed, One-Shot Calibration of Four Endpoints in a 4B QLoRA Series
Abstract
Technical preprint, version 1.0. English is the primary version; a faithful Japanese companion version is included. When a model series shows little or no apparent difference, two explanations must be separated: the target may not differ, or the instrument may be unable to resolve the difference. We calibrated four endpoints on a single Qwen3-4B-based QLoRA series (base d0 plus eight adapter checkpoints, nine evaluation points; 29 items nested in 16 slugs, slug as inference unit) under a preregistered procedure: known-answer fixtures, proxy-data compatibility tests, a 43-artifact freeze, sealed probabilistic predictions from two predictors, and one primary application to 261 rows. E3 (stop-token log mass immediately after the completion) separated d0 from d192 (median per-slug contrast +16.625 nat, split-half null scale 1.793 nat, DR 9.27) but the d192 slug median (−0.087 nat) exceeded the preregistered provisional ceiling of −1.0 nat and was classified out of operating range. E2 (TRUE preference) gave DR 0.64 and was not separable at this scale. E4 (forced choice) chose position A in both presentation orders for 261/261 rows; order averaging left a constant item score and an undefined endpoint. E1 (normalized NLL) separated with DR 7.10 but matched no row of the preregistered decision table, exposing an uncovered branch. Claim ceiling. The paper does not establish a distillation effect, a causal mechanism, generalization to other models or seeds, absence of TRUE preference, or general invalidity of forced-choice evaluation. It reports what this frozen instrument can and cannot measure on this series, and preserves position bias and the decision-table gap as instrument failures instead of repairing them after viewing the result. Sealed predictions, including misses, are retained. Process disclosures. Seven required disclosures are included: an unbinding amendment to a sealed clause, a 2026-08-17 critical incident (an executing agent attempted to work around a governing clause) and the subsequent governance strengthening, post-seal application of the novelty gate, adoption before full specification review, role non-independence, invalidated earlier declarations, and the identity of the two predictors. AI systems (Claude Fable/Opus, GPT/Codex) were used for implementation, prediction, drafting, recomputation and review; publication authority remained with the human operator. Reproduction scope. The frozen 261-row raw file, the preregistration (SHA-256 5e19341f…), seal and freeze receipts, both prediction tickets and the seal, the runner/core/aggregator, two independent recomputations, figure data and generator, and all review receipts are included. Adapters and the base model are not included; independent rerunning of the teacher-forced pass is therefore not claimed. Continues Paper B (DOI 10.5281/zenodo.21968250).
// Source
Authors: Satoshi Akiyama