Interventional Memory: A Pre-Registered Randomised Crossover Measuring Whether Agent Memory Changes Agent Behaviour
Abstract
VERSION 1.10 (2026-08-17) — what changed, and why there is a new version at all. This version corrects the deposited record rather than extending it. Version 1.9 shipped with four statements about how far the treatment reaches that had been computed under an earlier dose band and were never recomputed when the band was widened on 2026-08-16. One of them said that no locked dose reaches the brief’s eight primary slots; at the top dose it does. They are corrected in place, with the superseded text quoted beside each correction rather than deleted. Three further changes are substantive. The restriction of the boost to the two coverage slots is now a registered property of where the boost is applied, instead of an arithmetic consequence that a change of band could silently repeal. The scope rule, which excluded high-stakes sessions from treatment, is restated as an exclusion applied to the outcome rather than to the brief — a brief is composed at session start and the allowlist classifies actions that do not exist yet — and the ethics appendix is corrected to the weaker claim that is the true one: no high-consequence action drives a result, not that none followed a boosted brief. And claims_check.py is added, which recomputes every band-dependent quantity from the frozen constants and fails on divergence, because prose stating a computed result is a cache with no invalidation. No estimate, sample size, hypothesis or analysis rule changed. Version 1.9 remains retrievable under the same concept DOI, and the adversarial round that produced these corrections — including that three of four reviewer invocations returned no provider receipt — is recorded in the registration itself. THIS DEPOSIT IS A PRE-REGISTRATION: a study design, hypotheses and analysis plan fixed and published BEFORE the study runs. No outcome data exists at the time of deposit and no randomised epoch has been executed. Zenodo offers no dedicated resource type for pre-registrations, so this record is typed as a preprint and labelled here instead. Retrieval metrics score representation, not decision. This study measures whether the composition of an agent's memory changes what the agent does, using a fleet-wide randomised crossover on live production traffic rather than a curated benchmark. Epochs of 24 h are assigned to arms by a public randomness beacon whose round is declared before it exists; outcomes are adjudicated blind by a frozen multi-model panel; the primary outcome is repeated-failure density per session-hour. N = 234 epochs, powered for relative effects of 30% or more, sized on the upper confidence limit of the intra-cluster correlation (ICC 0.0985, 95% CI [0.0570; 0.1814]) rather than on its point estimate, and with a design effect that accounts for unequal cluster sizes (cv-squared = 0.3833, a locked input measured over the same pilot window as every other sizing parameter). The treatment writes one memory chunk per adjudicated-failure episode, in both arms, and differs only in a serving-time salience boost, so the contrast isolates weighting from creation. Measurement performed before registration establishes what the boost can reach. A chunk written from a failure episode enters a brief only through two coverage slots, and its reach depends on the severity it carries and on its age. Under the locked designation rule — the treatment boosts one chunk per failure signature present in the brief, the easiest of that signature to reach — the lowest dose in the band reaches 58% of opportunities and bounds the achievable effect at 60%, against a minimum detectable effect of 30%; the top dose reaches all of them, the modal S1 failure included. An earlier specification that boosted every match reached 19.8% and bounded the effect below the MDE, which would have made the primary hypothesis untestable; that is recorded in the deposit rather than removed. The three doses are pooled as one treatment arm for the primary contrast, so the dose-response comparison between them runs at 39 epochs per dose and is powered only for effects near 49.5% — a null there is reported as not detectable, never as evidence against the mechanism. These bounds are declared here rather than discovered later. VERIFIABILITY. The deposit carries the SHA-256 manifests of the frozen episode corpus, of the locked extractor and of the adjudication prompt; the three sampling seeds, each derived from a drand beacon round that was declared in a public commit before that round existed; and every analysis script, all deterministic and standard-library only. Anyone can recompute the sizing grid, the carry-over bound, the p95 winsorisation points and the intra-cluster correlation from the published inputs. WHAT IS DELIBERATELY ABSENT. No episode content is deposited. The adjudicated corpus contains real production work and stays out; this record points to it by hash. The only JSON Lines file included, positive-control.jsonl, holds six synthetic control episodes written by hand. All parameters were fixed on a historical corpus with no arm assignment, before any randomised epoch existed. Companion repository: https://github.com/totobusnello/memoria-nox (paper2-interventional/). WHAT CHANGED IN VERSION 1.11. An error correction, published before any epoch was randomised, before any arm was assigned and before the public beacon round that seeds the assignment was drawn. The design effect had been computed with the equal-cluster-size formula 1 + (m-bar - 1) * rho, while epoch sizes in this design run from 1 to 115 sessions. The applicable form is 1 + ((cv-squared + 1) * m-bar - 1) * rho, and the equal-size formula understates it, hence understates the variance, hence under-sizes the study — the one direction the registration's own sizing lock exists to forbid. N therefore moves from 174 to 234 and the allocation from 87 control / 29 per dose to 117 / 39. The minimum detectable effect stays at 30% and is not re-argued; the rate parameters, the ICC and its interval, m-bar, the carry-over constant, the dose band and every reachability number are untouched, because cv-squared enters the variance and not the rate. The dose-response contrast does not become stronger: K rose from 29 to 39 by nearly the factor the design effect rose, so its own detectable effect is 49.5% against 49.4% before. Version 1.9 remains retrievable under the same concept DOI, and both scripts still reproduce the superseded numbers on request. Version 1.10 was written and never deposited. This version also registers the coverage formula whose threshold had been locked since 2026-07-29 without the quantity ever being defined, and a carry-over channel through depletion of the never-served candidate pool that the snapshot argument does not bound.
// Source
Authors: Luiz Antonio Busnello