AI & Computingpreprint2026-08-09

Safety-Constrained Architecture Development: Preregistered Causal Experiments and Bounded Scientific Ledgers for AI Systems

Open access0 citations

Abstract

Exploratory AI architecture research gives the same team unusually broad control over benchmark construction, implementation, metric selection, threshold choice, reruns, and interpretation. That concentration of discretion creates opportunities for outcome-driven design changes, selective reporting, test-set leakage, overclaiming, and thedisappearance of failed directions. We present a repository-native framework for safety-constrained architecture development intended to make those failure modes visible and costly before external peer review is available. The implemented framework combines six core mechanisms: preregistration-first commit chronology with auditable ancestry; blind-evaluation gates with explicit unlock criteria and access counters; adjacent-contrast causal ladders that change one component at a time; bounded scientific ledgers that store claims together with their tested limits; a typed failure protocol that preserves negative results and resurrection conditions; and canonical numerical identities paired with executable decision invariants.We illustrate those six mechanisms with a three-experiment development-only sequence from Program K, a private program studying consequence-aware epistemic mechanisms. A matched-readout intervention improved several utility measures but breached a preregistered critical-safety allowance and was not accepted for operational use.Subsequent single-variable interventions localized the safety failure first to a sufficiency channel and then to cross-field applicability within that channel, while again rejecting the diagnostic intervention as an operational architecture because it created excessive false gaps. The sequence also exposed a provenance ambiguity in themiddle experiment: its scientific rules were fixed in an authoritative repository issue before evaluation, but its branch history did not preserve a separately auditable preregistration-only commit. That deficiency was recorded and led to a stricter chronology rule that the final experiment then satisfied. After the case sequence, an independent read-only evaluation identified additional program-level weaknesses. On 31 July 2026 the roadmap prospectively adopted a family-termination rule, a strong contemporary-baseline checkpoint, a preregistered naturalistic-data probe, external adversarial review, and explicit safety-coverage and cost reporting. These amendments did not govern the historical case study and are presented as review-driven hardening, not retroactive evidence of prior practice. We present the overall record as evidence of auditability, bounded causal attribution, and process correction, not as evidence that the governance framework improves scientific outcomes or that Program K generalizes beyond its synthetic benchmark.This record contains the preprint manuscript. The accompanying sanitized evidence bundle is available at https://doi.org/10.5281/zenodo.21726365.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-09

Authors: Vlad Zotta