Society & Economicspreprint2026-08-23

The Catalogue Is the Instrument: how the choice of sign inventory moves the statistics of an undeciphered script

Open access0 citations

Abstract

Statistical arguments about undeciphered scripts — whether a sign system "behaves like language" — are computed on a transliteration, and a transliteration presupposes a catalogue of signs. For rongorongo the published catalogues disagree by an order of magnitude (about 50 to about 600 basic signs). Pozdniakov and Pozdniakov (2007) said that statistics "will differ markedly depending upon the inventory of glyphs chosen"; the point has not, to our knowledge, been quantified. We quantify it on the CEIPP corpus (25 objects, 14,812 coded glyphs) using the normalised conditional entropy of Rao et al. (2010), on whose [0,1] scale the language-versus-non-language comparisons are drawn. Three published readings of the same inscriptions move the statistic by 0.318 of that scale; the corpus's distance from its own unigram-matched shuffle — the structure the statistic is meant to detect — is 0.077. The choice of catalogue therefore moves the instrument about four times further than the signal (3.7x after Miller-Madow correction). A sweep over catalogue sizes from 25 to 633 under four merge rules shows the effect is mostly size: signal grows monotonically with the inventory, coverage of the bigram table collapses beyond about 100-125 signs, and the field's own ~125-sign catalogue sits where the two criteria cross. Composition matters at fixed size, but differently for different questions: it moves conditional entropy by 0.004 when the frequent core is held fixed and by 0.06-0.14 when the merge rule may touch it, while against the scribes' own substitutions in parallel passages the published catalogue absorbs 40% of variants where the best mechanical rule of the same size absorbs 18% and chance 0.6%. Finally, we place three real syllabaries — Rapa Nui and Maori text syllabified, and Linear B — on the same instrument at matched size and length. At 50-55 signs, rongorongo can be merged so that its normalised conditional entropy is indistinguishable from syllabified Rapa Nui (0.707 vs 0.695), confirming that a small catalogue produces the syllabary's number; but its distance from its own shuffle is 2.4-3.6 times smaller than that of every syllabary tested even under the best rongorongo catalogue at each size, and up to ten times smaller under typical ones; the gap is unchanged when run lengths are matched and is flat in sample size. Corrupting 5% of glyph tokens moves the statistic by 0.011, so transliteration error cannot account for either gap. We conclude that levels of these statistics are properties of the catalogue and the sample; only differences — between readings, and between a corpus and its own null — travel. Version 3 (23 August 2026): Appendix A is rewritten and one of its conclusions is WITHDRAWN. Versions 1 and 2 reported rongorongo's repetition rates against a shuffle of the pooled 25-object corpus and concluded that a plain syllabic reading of a Polynesian language was ruled out. Against a null matched to the way the reference languages were measured - one shuffle per text - AA falls from 1.6-2.4 to 1.12, AAA from 5-15 to 1.68 and ABAB from 30-100 to 20.7, and Linear B, an actual syllabary, exceeds rongorongo on AAA. The elimination argument does not survive and is withdrawn; what does survive is stated in its place. Section 2 also notes two CEIPP codes that are not signs and were counted as signs in two of the three readings; correcting that moves the headline ratio from 4.14 to 4.28, so the published figure was conservative.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-23

Authors: Ilpo Väätäinen