Same Content, Wider Track: Empirical Calibration of Friction Theory on LLM Substrate
Abstract
The same 25 facts, fine-tuned with four paraphrase templates instead of one, jump from 38% to 94% paraphrase-robust recall (same substrate, same optimizer, same number of facts, same epochs). Varied friction widens the encoding track. Paper 4 is the empirical-calibration companion to Paper 6 (matched-friction-under-hysteresis), tested on LLM substrate (Qwen2.5-7B-Instruct + LoRA fine-tuning). Abstract. Friction Theory (FT; Lund 2026b, Paper 1) is a mechanism-level framework for bounded probabilistic computation under cost in systems satisfying the race axioms, standing on sequential-sampling and resource-rational accounts; its cross-substrate reading is a working hypothesis proposed in Paper 1. The matched-friction-under-hysteresis formulation (Paper 6) predicts an inverted-U between substrate-level friction intensity and encoding outcome on any axis where friction scales monotonically. Behavioural Friction Theory (BFT; Lund 2026a, Paper 0) applies the framework to mobile, mortal, metabolism-constrained organisms; the nesting of BFT within FT was proposed in Paper 1 and remains a working hypothesis. Paper 4 is the empirical-calibration companion on LLM substrate. Methods. Eight friction-intensity axes tested via fine-tuning Qwen2.5-7B-Instruct (4-bit NF4 LoRA, r=16, α=32) on fictive "Zorbetik" facts designed to eliminate pretraining priors, plus cross-substrate-paradigm tests on Mistral-7B-Instruct-v0.3, Llama-3.1-8B-Instruct, Llama-3.1-8B-base, and Qwen2.5-3B-Instruct via FT trajectory and ICL-inference paradigms. Headline results. Four of five intra-session axes produce genuine inverted-U parabels: task-friction (v10c, R_LLM ratio ordering deep_reason 0.66 < passive 0.75 < surface_sort 0.95, cross-family replicated); chunking density (R² = 0.89); learning rate (catastrophic cliff σ = 0.019 log-units); sampling temperature (R² = 0.985, peak T ≈ 0.4–0.5). The fifth axis (violation magnitude, v11c) produces a framework-narrowing null on behavioural accuracy; v12-consolidation differential preservation under orchestrated replay confirms Level-0 encoding-differentiation via distinctiveness-confidence rather than surprise-magnitude-proportionality. The Bjork-spacing bonusparabel (v4c/v4d) produces a formally-undecided null traced to architectural absence of internal between-session replay. Distribution-shape extension (§4.7, NEW). A 7th and 8th axis sweep training-distribution shape rather than friction-intensity. The v12b paraphrase-augmentation experiment provides the cleanest within-content demonstration of distribution-shape as substrate-friction-parameter: same 25 facts trained with 1 paraphrase-template yield 38% paraphrase-robust memorisation; the same facts trained with 4 paraphrase-templates yield 94% (lift +56 pp under matched substrate, optimizer, total facts, total epochs). This is the empirical anchor for "varied friction = wider track" inherited from the human desirable-difficulties literature (Schmidt & Bjork 1992; Rohrer & Pashler 2007). Three Paper-6 refinements forced by Paper 4 data. σ is strongly axis-dependent (~13× span between LR and temperature; Paper 6 §4.11 axis-specific kinetics). The Michaelis-Menten × sigmoid-gate formula requires a baseline-offset extension for axes on already-trained substrate (Paper 6 §4.5). The v11c null + v12 differential preservation force a four-way consolidation taxonomy (Paper 6 §4.10): functional consolidation (BFT, biological), orchestrated mechanical analog (FT, LLM + replay), pure mechanical substrate (FT, LLM without replay), tool-augmented mechanical consolidation (FT, LLM + retrieval, untested but framework-predicted). Distribution-shape cross-substrate matrix. HRP-3M direction (deep > passive ≈ surface in correctness or first-token CR distribution) is confirmed on 5 of 6 substrate × paradigm pairings spanning two model families (Qwen, Llama), three sizes (3B/7B/8B), and two paradigms (FT trajectory + ICL inference). The 6th pairing is ceiling-masked under the pre-flight test-informativeness criterion, consistent with that criterion. Two substrate-distinguishing friction-type axes (§11.6). Race-level friction including reactance recurs across the substrate classes tested (Paper 14, Paper 2B, Paper 5 report the anchors); whether it is substrate-general remains an open prediction. Two axes distinguish LLM substrate from biological substrate: self-continuity (confirmed substrate-absent via §4.8 v10g), and surprise (biology-specific via §4.3 v11c null + §11.5 "Dunning-Kruger without Mount Stupid" theoretical extension). Scope and status. Paper 4 is a pilot-scale preprint timestamping preliminary empirical observations on LLM substrate. Per-condition n in the range 4–30; mixed-model, mixed-paradigm portfolio. Individual findings should be read as exploratory pilot results that constrain framework development, not as definitive cross-substrate calibrations. A planned v2 revision will scale per-condition n on the core axes; see §10.1. A pilot scale-up attempt (Together SFT, n=100/condition at both LoRA r=16 and r=64) produced 0% accuracy — substrate-saturation finding documented in §10.1 as scope-condition rather than framework-disconfirmation; v2 revision must combine n-expansion with paraphrase-augmentation per §4.7.3. Companion papers in the Friction Theory series: Paper 0 (BFT master): 10.5281/zenodo.19462500 Paper 1 (Friction Theory substrate): 10.5281/zenodo.20012655 Paper 2B (ICL/FT substrate-mechanism): 10.5281/zenodo.20145219 Paper 4B (Expertise reversal spinoff): 10.5281/zenodo.20059862 Paper 6 core (Matched friction under hysteresis): 10.5281/zenodo.20059863 Paper 6B (Effort-value attribution): 10.5281/zenodo.20309000 (cite-letter; in preparation) Paper 6C (Dynamic commitment timing): 10.5281/zenodo.20309001 (cite-letter; in preparation) Series position. Paper 4 in the Friction Theory paper-series. v2 (August 2026) — scope and reporting revision. Claim scope is tightened to what the axes support: the headline generalisation is restricted to the intensity and dissociation axes, per-axis outcomes are reported against the study's own replication map rather than summarised upward, and one instance withdrawn on instrument grounds is relabelled as a v1 instrument record. The framework is positioned with priority credited to the prior sequential-sampling and resource-rational accounts, and its cross-substrate reading is carried as a working hypothesis rather than a premise. Editorial pass. Earlier versions remain in the version history.
// Source
Authors: Tomas Pødenphant Lund
Institutions: Aarhus University