Society & Economicspreprint2026-08-23

The Catalogue Is the Instrument: how the choice of sign inventory moves the statistics of an undeciphered script

Open access0 citations

Abstract

Statistical arguments about undeciphered scripts — whether a sign system "behaves like language" — are computed on a transliteration, and a transliteration presupposes a catalogue of signs. For rongorongo the published catalogues disagree by an order of magnitude (about 50 to about 600 basic signs). Pozdniakov and Pozdniakov (2007) said that statistics "will differ markedly depending upon the inventory of glyphs chosen"; the point has not, to our knowledge, been quantified. We quantify it on the CEIPP corpus (25 objects, 14,812 coded glyphs) using the normalised conditional entropy of Rao et al. (2010), on whose [0,1] scale the language-versus-non-language comparisons are drawn. Three published readings of the same inscriptions move the statistic by 0.318 of that scale; the corpus's distance from its own unigram-matched shuffle — the structure the statistic is meant to detect — is 0.077. The choice of catalogue therefore moves the instrument about four times further than the signal (3.7x after Miller-Madow correction). A sweep over catalogue sizes from 25 to 633 under four merge rules shows the effect is mostly size: signal grows monotonically with the inventory, coverage of the bigram table collapses beyond about 100-125 signs, and the field's own ~125-sign catalogue sits where the two criteria cross. Composition matters at fixed size, but differently for different questions: it moves conditional entropy by 0.004 when the frequent core is held fixed and by 0.06-0.14 when the merge rule may touch it, while against the scribes' own substitutions in parallel passages the published catalogue absorbs 40% of variants where the best mechanical rule of the same size absorbs 18% and chance 0.6%. Finally, we place three real syllabaries — Rapa Nui and Maori text syllabified, and Linear B — on the same instrument at matched size and length. At 50-55 signs, rongorongo can be merged so that its normalised conditional entropy is indistinguishable from syllabified Rapa Nui (0.707 vs 0.695), confirming that a small catalogue produces the syllabary's number; but its distance from its own shuffle is 2.4-3.6 times smaller than that of every syllabary tested even under the best rongorongo catalogue at each size, and up to ten times smaller under typical ones; the gap is unchanged when run lengths are matched and is flat in sample size. Corrupting 5% of glyph tokens moves the statistic by 0.011, so transliteration error cannot account for either gap. We conclude that levels of these statistics are properties of the catalogue and the sample; only differences — between readings, and between a corpus and its own null — travel. Version 2 (23 August 2026): section 7 gains four collections of native Rapa Nui prose supplied by E. Korovina, which confirm the syllabary gap and show that one dissenting collection dissents because of its transcription rather than its language. Version 1 gave rongorongo's period-two repetition as 30-100 times expectation without stating the same figure for the reference languages, which are at 2.5-7.7 rather than near 1; the contrast is fourfold to sixfold, and Appendix A now says so.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-23

Authors: Ilpo Väätäinen