rmt-llm-research — Random Matrix Theory meets Large Language Models: spectral analysis of activations and the mathematical inevitability of hallucinations
Abstract
A research repository by Iskhak Hamzatovich Isaev presenting a comprehensive program applying Random Matrix Theory (RMT) to the analysis of large language models (LLMs). Thread 1 — Spectral statistics of LLM activations. Covariance matrices of GPT-2 hidden-state activations exhibit fundamentally different spectral properties in factual recall vs creative generation mode. The BBP phase transition, Marchenko-Pastur bounds, and Tracy-Widom fluctuations provide quantitative markers for cognitive mode detection. Thread 2 — Mathematical proof of inevitable hallucinations. Ten independent mathematical paths converge on a single critical token count N_crit: complex phase, operator dynamics, GUE spectral statistics, thermodynamic entropy, Lévy-Langevin fractional dynamics, EP-surfaces, Tracy-Widom, quantum channel degradation, NHSE (winding number), and Caputo fractional-time memory. Thread 3 — The utility trap: RLHF and entropic collapse. RLHF optimization creates an artificial drift in the Fokker-Planck equation, forcing the model to «lie beautifully once» rather than risk a self-correction cycle. With Caputo memory parameter β ≈ 0.5, the mean collapse time scales as ⟨T_crit⟩ ∝ (μ_eff)-2 — even small RLHF pressure quadratically accelerates hallucination onset. The repository ships 12 DOCX monographs (EN/RU), 1 PDF paper, a GitHub Pages documentation site, an interactive in-browser demo, a Python verification package (rmt_llm) with 70+ pytest tests, a Julia verification package (RMTLLMVerify), and a Jupyter verification notebook.
// Source
Authors: Iskhak Hamzatovich Isaev
Institutions: Independent University of Moscow