A Field-Theoretic Reading of Transformer Probing Structure: QKV Three-Channel Division × q Two-Layer × Final-Layer Immunity
Abstract
The query-key-value (QKV) three-channel architecture is the core of the attention mechanism, yet a unified reading of the functional division among these channels within the probing structure remains lacking. Using a field-theoretic framework (materials as density-clustering fields ρ=m·p; the information two-axis = density axis [where to probe] × adsorption axis [how long to persist]), this paper provides a computable empirical account of QKV division across the Qwen2.5/Qwen3 series (0.5B–7B, 7 models). Core results: (1) free parameters = the active driver of diffuse probing (H2, four layers of evidence closed); (2) q probing = two-layer structure — a discriminative core of 48 dims (1.34%, directional rather than dimensional, globally shared across heads × graded singular-value envelope) × a generative-matching channel of 768+ dims (incompressible); q redundancy is 98.7%, yet redundancy is compressible while the cooperative ontology is not; (3) final-2-layer immunity — the functional core is ≤8 dims, mechanism of downstream residual-stream absorption; (4) k/v + MLP memory is incompressible — threshold = k/v projected output dimension; (5) q-spectrum tail = broad-spectrum diffuse probing direction — top-96 carries only 23.6% of the energy while the tail carries 76.4%; the tail bears generative distributional detail and discriminative candidate comparison; amplifying it collapses the model; tail energy grows monotonically with scale (r=0.969 across families); (6) structural destruction > signal nullification — in the partial-destruction regime M1 is significantly more harmful than M5. All values are machine-verified; scripts are archived as a reproducibility package; observation/inference are separated and architectural specificity is explicitly annotated. Chinese version is the companion translation.
// Source
Authors: Chao Qin
Institutions: Minzu Normal University of Xingyi