pi-canon: The Write Desk
Abstract
The first two pi-canon papers measured recall: whether a canonical project memory can be surfaced when it is needed, and what that recall costs. Both took for granted that what the store holds is true. This third report tests that assumption and finds it does not hold. The instrument is a lineage of session records whose facts are revised, reversed, and retired across sessions, written into a real store by a real agent through the shipped tool, then read back by a fresh agent whose answers are graded per slot against a hidden oracle. The endpoint is the number of superseded values still standing in records whose contract is to state what is true now. The main result is a direction without a magnitude. Across three controlled captures on two models, writers repeatedly left superseded values in place, and a condition in which the tool spoke at the write boundary ended lower on that endpoint. Across 24 capture-lineage comparisons over the same eight constructed histories the feedback condition ended lower in 20, tied in 2, and higher in 2. The size did not survive: a counterbalanced repeat that held model, fixture, prompts, and arm tools, but reversed arm order in half the lineages and differed in the uniform state of two pinned package files, moved a treated condition far enough to erase most of the difference the first capture reported. One lineage of eight supplied 10 of that capture's 14-point pooled contrast. Four of five predictions registered mid-capture and before grading failed. The magnitude is withdrawn rather than qualified, and with one realization per cell the design cannot estimate run-to-run variation, let alone rank it against the arm contrast. Read the contrasts as differences between conditions, not as a treatment effect identified on the message. A separate offline retrieval benchmark on two real project stores put the shipped lexical ranker ahead of cosine similarity over local embeddings in eight conditions of eight, by 15.8 to 37.0 points. That benchmark had been reading its corpus live; it is frozen here, and the report includes the measured cost of two defects found in the freezing itself by external review. The supplement carries every results document, the per-lineage endpoint CSV behind every direction count, the graded reports, the digest-pinned capture contracts, the fixture generators, both arm tools with the diff between them, and scripts that recompute the paper's quantitative claims from the artifacts rather than restating them. The frozen retrieval corpus is two real private project-memory stores; its digest is published here and its contents are not. pi-canon is MIT licensed and is on npm. This record is a preprint and has not been peer reviewed; it carries eleven rounds of adversarial model review, and the errata for claims corrected or withdrawn across those rounds ship with the supplement.
// Source
Authors: Shane Conner