AI & Computingpreprint2026-09-09

VARZIN V2 — Recoverable but Not Composable: Representation Recovery and Systematic Composition in a Frozen Language Model

Open access0 citations

Abstract

This research report presents the integrated Phase 0–2 results of VARZIN V2, an experimental study of the relationship between structural representation recovery and systematic compositional generalization in a frozen language model. The study asks whether a structural variable that is strongly encoded and recoverable from a frozen language model's internal representations is consequently usable as a computational substrate for systematic composition on unseen combinations. Phase 0 — Dataset construction and closure. A controlled synthetic morphological-position dataset was constructed and independently audited for multiple forms of leakage, including token-bag, character-bag, boundary/BPE, and machine-learning leakage. The final dataset contains 10,800 records, 12 structural positions, four synthetic morphemes, 30 contexts, 24 training contexts, and 6 held-out contexts. After verification with the actual Qwen tokenizer, the dataset was cryptographically frozen and used unchanged throughout the subsequent experiments. Phase 1 — Representation recovery. The frozen Qwen2.5-1.5B-Instruct model was evaluated at all 29 hidden-state indices (embedding plus layers 1–28). The preregistered primary representation, surface-span mean, yielded extremely high held-out-context linear-probe performance, with a mean balanced accuracy of approximately 0.9666 across the 29 hidden-state indices and 25 of 29 indices reaching balanced accuracy of 1.0000. These results demonstrate that the tested synthetic structural variable is highly linearly decodable from the frozen model's hidden representations under held-out-context evaluation. Phase 2 — Composition generalization. The study then tested whether the recovered representations support systematic generalization of the preregistered operation k = (i + j) mod 12 to unseen ordered pairs. The experiment evaluated 144 ordered pairs, with 120 training pairs and 24 fixed holdout pairs, across all 29 hidden-state indices, two fixed composer architectures, four label conditions, and 10 preregistered training seeds. The primary evaluation cell was training contexts × holdout pairs. The composer received only the concatenated representations [z_i ; z_j] and did not receive the underlying structural indices directly. The preregistered positive criterion required the TRUE condition to outperform each control by at least 0.15 balanced accuracy and to pass paired exact sign-flip tests at p < 0.01. No layer/composer combination satisfied the complete criterion: 0/58 PASS, 0/58 AMBIGUOUS, and 58/58 FAIL. The strongest observed TRUE-versus-control effect was +0.082 balanced accuracy, for the fixed MLP composer at Layer 5 versus the RANDOM condition, with an exact sign-flip p-value of 0.1719. Diagnostic analyses indicate that this result was not simply caused by constant-output classifier collapse. In the primary evaluation cell, 93.0% of linear-composer rows and 83.2% of MLP-composer rows used the full 12-class output space. The combined result therefore provides a controlled empirical instance in which strong representation recovery was observed while the corresponding preregistered test of systematic composition over unseen combinations failed. The study does not establish that algebraic structure is absent from the model, that the model cannot perform the tested operation by another mechanism, or that language models generally lack compositional computation. It also does not establish a universal principle that decodability never implies composability. The principal methodological implication is narrower: strong decodability of a structural variable should not, by itself, be treated as evidence that the same representation functions as a computational substrate for systematic compositional generalization. An explicitly exploratory post-hoc observation concerns the Latin-square structure shared by the TRUE, SHUFFLED, and WRONG_OP conditions. This observation was not preregistered and is therefore treated only as hypothesis-generating evidence for future work. The report includes the preregistered experimental design, final decision rule, implementation amendment, dataset-freeze information, cryptographic artifact hashes, reproducibility information, limitations, and proposed future experiments. Keywords: frozen language models; representation learning; probing; compositional generalization; mechanistic interpretability; algebraic structure; synthetic language; morphology; systematic generalization; reproducibility; Qwen2.5-1.5B-Instruct; VARZIN V2

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-09-09

Authors: Nirouyar Reza

Institutions: Varian Medical Systems (Switzerland)