Structured Documentary Access and Cross-Session Reconstruction Fidelity in Large Language Models: A Controlled Replication Across Three Agents and Three Domains
Abstract
Large language models operating without persistent memory must reconstruct project context from whatever is available at query time. We report a controlled comparison of reconstruction fidelity under three conditions: no documentary access (SIN), access to the identical corpus with computed classification metadata removed but session structure intact (STRIPPED), and access to a structured documentary corpus (Codex Bookshelf, CB). Using a closed codebook of 75 functional recovery units (UFR) scored across three independent domains and evaluated by three different language model agents (Claude, Codex, Kimi), we find reconstruction fidelity of 24.0% under SIN, 84.7-94.7% under STRIPPED, and 97.3-98.7% under CB, a monotonic pattern that replicates consistently across all three agents and all three domains. The STRIPPED condition was scored by an evaluator blind to both agent identity and experimental condition, a stricter coding procedure than the one applied to SIN/CB. We report the full coding methodology, inter-round audit corrections, and an explicit discussion of what this design does and does not establish: STRIPPED narrows, but does not close, the gap between access and organization. It removes derived classification fields while leaving session boundaries and occurrence-evidence intact, and is therefore not equivalent to an equal-volume, structurally disorganized corpus. We describe the remaining, narrower question as ongoing work, in keeping with the pre-registered principle that no architectural claim should precede the evidence that supports it. This version (v2) adds the STRIPPED condition, its blind-coding methodology and verification incident, an expanded results table, and expanded limitations relative to v1.
// Source
Authors: Claudio Martín Sosa
Institutions: ASTER