Word-Order Information, Page-Anchored Repetition, and a Data-Estimated Generative Model for the Voynich Manuscript
Abstract
The Voynich Manuscript (Beinecke MS 408) combines word-level statistics that are ordinary for a natural language with character-level statistics that are not. We report a layered measurement battery applied to the Zandbergen-Landini ZL3b transliteration (34,103 running-text tokens, 7,417 types) alongside seven medieval reference corpora processed with identical code, quantifying an explicit noise floor for every information-theoretic estimator used. Four primary results carry the argument: Second-order conditional character entropy is h2 = 2.17 bits at a matched letter budget, while all seven reference corpora fall between 3.02 and 3.52 bits. A controlled simulation of Latin scribal abbreviation moves h2 upward (to 3.40 bits) and cannot account for the deficit. A two-part minimum description length (MDL) criterion places the compression optimum at an inventory of 426 sub-word units of mean length 3.41 characters. These units sit at syllable scale: unit entropy is 7.86 bits compared to 7.88 bits for 14th-century Italian syllables and 7.94 bits for Latin. On transition graphs, the manuscript lies closer to Italian (normalized Laplacian spectral distance 1.22) than Latin (1.74), but further from either than the two reference languages lie from each other (0.56). Excess mutual information between tokens is Delta I(1) = 0.243 bits. From lag 2 onward, it is statistically indistinguishable from an estimator noise floor of 0.011 bits. In contrast, Latin and Middle English retain 0.128 and 0.214 bits over lags 2 to 10. The mid-range syntactic word-order information is absent. At matched token lag, the excess of near-identical word repetition falls by a factor of 1.41 across folio boundaries (z = 6.4 against a circular-shift null) and by 1.29 across paragraph boundaries within a folio. This locates the word reuse window on the physical parchment page rather than in the sequential token stream. A fixed-rotation volvelle is excluded by a permutation periodogram (p = 0.31). A generative model estimated entirely from the manuscript and evaluated on its own page skeleton reproduces 9 of 14 target statistics (modal count over 100 seeds). Its systematic residuals indicate that real page binding is stronger and real long-range mutual information is weaker than the model produces. Repeating principal measurements on three additional transliterations (including the v101 alphabet) confirms that character entropy remains in the 2.13 to 2.50 bits range and the word-order decay remains unchanged. These measurements do not exclude verbose homophonic ciphers, which destroy word-order information by construction while preserving plaintext messages. All analysis scripts, result files, and reproduction code are archived alongside this paper.
// Source
Authors: Lloyd Stellar