Interventional Memory: A Pre-Registered Randomised Crossover Measuring Whether Agent Memory Changes Agent Behaviour
Abstract
VERSION 1.12 (2026-08-26) — the amendment is descriptive, and it says so. This version adds AMENDMENT-v1.12.md, which does three things and refuses a fourth. It declares facts about the serving mechanism, the corpus and a pilot series; it retracts twenty-eight earlier claims, each with the measurement that replaces it; and it names the defects that must be fixed before the confirmatory study starts. It does not specify an estimand, does not fix N, and does not establish a causal effect. Why it cannot. In 2,221 of 2,221 serving decisions logged between 2026-08-21T22:57Z and 2026-08-25T10:22Z, the mode was shadow and the brief actually served was the control arm. The treated arm was computed, logged and discarded. With no outcome observed under treatment, N and the assignment timestamp are not estimable from this series, and any estimand written now would be post-observational — the series was seen before the wording. The 2,221 records are therefore declared a descriptive pilot that does not enter confirmatory analysis. The estimand, estimator, variance and target population go into a separate prospective registration, later than this deposit and later than the fixes below, labelled as pilot-informed. The open defect, stated rather than repaired. The designation rule — which chunk of a signature group receives the boost — is not validly frozen. It consumes a constant, CUT_FRESH = 0.7342, whose referent this amendment retracts: 0.7342 is the coverage-slot cut, and no threshold of that kind is applied by the code that composes a brief. The registered tie-break (w_min, then created_at, then chunk_id) names a column that does not exist in the verdict table, so it is not implementable as written rather than merely unimplemented; and the code implements no tie-break at all, so ties fall to whatever row order SQLite returns. Measured at the window close, exact ties in w_min occur in 4 of the 7 multi-member signature groups, out of 19 groups. In those, the designated chunk came from incidental row order. The aggregates in this amendment are therefore reproducible — a frozen window and a guard script pin them — but not attributable to a deterministic rule. DECISION-designacao-2026-08-25.md sets out the replacement options, the requirements a replacement must meet, and a recommendation. Numbers are emitted, not asserted. An earlier draft of this amendment cited the pilot count as a snapshot: n = 2,256 “measured at 10:22Z”. The cron writes 28 records an hour, so within ninety minutes the live log held 2,263 — and the frozen window ending at 10:22Z actually holds 2,221, meaning label and measurement never corresponded. In an immutable deposit such a figure becomes false on its own, with no edit to trigger review. The fix is structural: the window is declared as a closed interval, every aggregate over it is emitted by pilot_window_stats.mjs, and the script accepts --assert-json and exits non-zero if the window stops reproducing PILOT-WINDOW-2026-08-25.json. The clipped log is deposited as p2-serving-WINDOW-2026-08-25.ndjson (2,221 records, integer chunk ids and timestamps only, no corpus content), so a third party can run the guard rather than take the prose on trust. Also corrected in this version. Section 2 now publishes what the λ estimate needs in order to be reconstructed: the target population of 1,305 episodes, the two stratum frames (48 by census, 1,257 sampled at 242), the realised sampling rate, and the seeded selection rule — so that (44 + 5.194215 × 11) / 1,305 = 0.077499 can be checked in full. An adversarial review had reconstructed the population as 1,261.4 from the published components and concluded the record was inconsistent; the reconstruction was wrong, because the Horvitz-Thompson weight divides the frame by the sample and not by the adjudicated subset, but the gap it exposed was real and documentary. The five serving modules the amendment transcribes or cites by line are now deposited verbatim (serving-*.ts, provenance in SERVING-CODE-MANIFEST.md), closing a defect the amendment had deferred: the code carrying the mechanism lived only on a private host. They are auditable, not executable standalone, since their imports are not deposited. claims_check.py gains a structural check it was missing: a one-line variant of the reachability script carried its own dose-band declaration that nothing was parsing. Adversarial review. Seven passages across four model families, every one with a verified receipt. Two invocations that produced no analysis are recorded as such, with the reason, so that no result is attributed to them. Claims raised in review and refuted by measurement are listed with their refutations, and so is the converse case: a claim that was refuted while the doubt behind it turned out to be correct, and cost two retractions. VERSION 1.10 (2026-08-17) — what changed, and why there is a new version at all. This version corrects the deposited record rather than extending it. Version 1.9 shipped with four statements about how far the treatment reaches that had been computed under an earlier dose band and were never recomputed when the band was widened on 2026-08-16. One of them said that no locked dose reaches the brief’s eight primary slots; at the top dose it does. They are corrected in place, with the superseded text quoted beside each correction rather than deleted. Three further changes are substantive. The restriction of the boost to the two coverage slots is now a registered property of where the boost is applied, instead of an arithmetic consequence that a change of band could silently repeal. The scope rule, which excluded high-stakes sessions from treatment, is restated as an exclusion applied to the outcome rather than to the brief — a brief is composed at session start and the allowlist classifies actions that do not exist yet — and the ethics appendix is corrected to the weaker claim that is the true one: no high-consequence action drives a result, not that none followed a boosted brief. And claims_check.py is added, which recomputes every band-dependent quantity from the frozen constants and fails on divergence, because prose stating a computed result is a cache with no invalidation. No estimate, sample size, hypothesis or analysis rule changed. Version 1.9 remains retrievable under the same concept DOI, and the adversarial round that produced these corrections — including that three of four reviewer invocations returned no provider receipt — is recorded in the registration itself. THIS DEPOSIT IS A PRE-REGISTRATION: a study design, hypotheses and analysis plan fixed and published BEFORE the study runs. No outcome data exists at the time of deposit and no randomised epoch has been executed. Zenodo offers no dedicated resource type for pre-registrations, so this record is typed as a preprint and labelled here instead. Retrieval metrics score representation, not decision. This study measures whether the composition of an agent's memory changes what the agent does, using a fleet-wide randomised crossover on live production traffic rather than a curated benchmark. Epochs of 24 h are assigned to arms by a public randomness beacon whose round is declared before it exists; outcomes are adjudicated blind by a frozen multi-model panel; the primary outcome is repeated-failure density per session-hour. N = 234 epochs, powered for relative effects of 30% or more, sized on the upper confidence limit of the intra-cluster correlation (ICC 0.0985, 95% CI [0.0570; 0.1814]) rather than on its point estimate, and with a design effect that accounts for unequal cluster sizes (cv-squared = 0.3833, a locked input measured over the same pilot window as every other sizing parameter). The treatment writes one memory chunk per adjudicated-failure episode, in both arms, and differs only in a serving-time salience boost, so the contrast isolates weighting from creation. Measurement performed before registration establishes what the boost can reach. A chunk written from a failure episode enters a brief only through two coverage slots, and its reach depends on the severity it carries and on its age. Under the locked designation rule — the treatment boosts one chunk per failure signature present in the brief, the easiest of that signature to reach — the lowest dose in the band reaches 58% of opportunities and bounds the achievable effect at 60%, against a minimum detectable effect of 30%; the top dose reaches all of them, the modal S1 failure included. An earlier specification that boosted every match reached 19.8% and bounded the effect below the MDE, which would have made the primary hypothesis untestable; that is recorded in the deposit rather than removed. The three doses are pooled as one treatment arm for the primary contrast, so the dose-response comparison between them runs at 39 epochs per dose and is powered only for effects near 49.5% — a null there is reported as not detectable, never as evidence against the mechanism. These bounds are declared here rather than discovered later. VERIFIABILITY. The deposit carries the SHA-256 manifests of the frozen episode corpus, of the locked extractor and of the adjudication prompt; the three sampling seeds, each derived from a drand beacon round that was declared in a public commit before that round existed; and every analysis script, all deterministic and standard-library only. Anyone can recompute the sizing grid, the carry-over bound, the p95 winsorisation points and the intra-cluster correlation from the published inputs. WHAT IS DELIBERATELY ABSENT. No episode content is deposited. The adjudicated corpus contains real production work and stays out; this record points to it by hash. The only JSON Lines file included, positive-control.jsonl, holds six synthetic control episodes written by hand. All parameters were fixed on a historical corpus with no arm assignment, before any randomised epoch existed. Companion repository: https://github.com/toto
// Source
Authors: Luiz Antonio Busnello