Who Validates the Validator? Instrument failure in the shape of the hypothesis — fourteen audit findings from a pre-registered measurement
Abstract
v5 (2026-08-08) — record integrity, power boundary, regime surfaces, finding register. Correction #30: the showcase builder wrote platform-default line endings while claiming determinism; output is now pinned to LF, and a new RECORD_MANIFEST.sha256 binds every file of this record — the showcase layer previously sat outside every gate. A Linux rebuild is now byte-identical to the Windows-built uploads on all eight showcase files, measured rather than assumed. Power boundary: the pre-registered design was sized for a base rate near 0.15; at the observed 0.80% it has no meaningful power for H2 — the null result is absence of evidence, not evidence of absence, now stated beside the number in both results documents. Measurement regimes: the report contract (fenced-json-results/1), protocol topology (orchestrator-worker-text-protocol/1) and tool-name surface (fixed-toolset-disclosed-names/1) join the error surface with named regime identifiers in docs/measurement_regime.json: the 0.80% is an estimate for prose-reporting regimes, not for agentic systems in general — under native structured tool calling the call is the claim and cannot diverge. Finding register: all thirty findings with date, channel, layer and phase in docs/finding_register.json; exposure units are not yet modeled and the recorded proxies are named as proxies. The registered primary outcome is unchanged: 10/1250 = 0.80%. v4 (2026-08-08) — boundary condition on the silent-substitution rate. Correction item #14 reported a silent-substitution rate of 18/91 = 19.8% without stating what the agent could SEE after a tool failed. The logging proxy does not merely observe: it produces the error representation the agent then reasons from — ERROR:<ExceptionType>, with no message, no traceback and no native exception propagated. The proxy therefore functioned as an interface intervention as well as a recorder, and the estimate is conditional on that representation; it should not be compared directly with measurements taken under native exceptions, structured framework error objects or full tracebacks unless error-surface policy is treated as an experimental factor. All 18 cases also arose on a single tool (text_stat). docs/error_surface.json names the regime (proxy-normalized-type-only/1) so a future figure is a different NAMED regime rather than an unexplained disagreement. The registered primary outcome is unaffected: 10/1250 = 0.80% — it asks whether a claimed call happened, which does not depend on how a failure was displayed. The papers are deliberately unchanged: dated texts with shipped PDFs and no reproducible build recipe. v3 (2026-08-07) — packaging-only correction. The parameterless release gate previously referenced the v1 archive name, so after v2 shipped it gated the earlier archive — and the parameterless invocation is the one the reproduction instructions describe as "one command". Default package discovery is now version-independent and fails closed on zero or multiple candidates. The correction protocol gains section 26 for this item, and sections 16–25 pointing to findings recorded in the instrument's portable form (dispatch-fidelity). No research data, task outputs, scores, statistical analyses, registered outcomes or conclusions changed. Gate logs for Windows and Linux verify the same archive hash. v2 (2026-08-07) — external finding #15. This version adds correction item #15 to the correction protocol: cross-run evidence splicing. File-level integrity proves that a file is what it claims to be; it proves nothing about whether two files belong together. Genuine, hash-valid artifacts drawn from different runs assemble into a package that every v1 check passes. The binding material was already in the package and had never been checked: the run manifest commits sha256(nonce) before execution, and the plaintext nonce is recoverable only from a genuine canary receipt in the tool log. Two new analyses run that check — analysis/check_run_binding.py (B1–B5; B3 proven for 125 of 129 run groups, four declared unprovable in docs/binding_unprovable.txt) and analysis/x2_splice.py (X2 injection class; manifest and tool-log splices detected at 100.000%, specificity 120/120). The registered primary outcome is unchanged: 10/1250 = 0.80%. The item concerns the structure of the record, not the numbers of the measurement; no artifact in the package proved to be spliced. The finding arrived after publication, on the open call in the Attack Map — the first to do so. Further findings will produce further versions; that is what this record is for. A pre-registered measurement of dispatch fidelity in multi-agent LLM orchestrations: whether a claim that a tool was called corresponds to an execution that actually happened. Ground truth comes from an out-of-band logging proxy with a per-run canary nonce; scoring is deterministic (MATCHED / FABRICATED / OMITTED). Registered primary outcome: 10 / 1250 = 0.80% fabricated dispatch claims [0.44–1.47%]. The pre-registered positive scaling hypothesis was not supported (Cochran–Armitage z = −1.106, two-sided p = 0.269). A preliminary v1.1 scorer had reported 54 / 1250 = 4.32% with a strongly significant positive trend; that series is retracted, and the defect that produced it took the shape of the expected result. The deposit documents fifteen audit findings — fourteen from the campaign and one that arrived after publication, from an external reader — and the five mechanisms that caught them. The scorer itself was validated by fault injection (18 classes; pooled sensitivity 169/169, specificity 131/131, plus 20 negative controls), including an inert injection check: an injected defect must be shown to be live, not merely well-formed. Verification. The package ships its own read-only, fail-closed verifier, a manifest, a full seal chain, and a release gate that extracts the archive into a clean room, runs every documented command, and requires that nothing in the shipped tree changes. Gate logs for Windows and Linux are included.
// Source
Authors: Zoltan Varga