Six Nameplates, Two Effective Votes: A Pre-Registered Local Replication of Correlated LLM Judge Errors, with a Deterministic Instrument for Measuring Panel Independence
Abstract
Kohli (2026, arXiv:2605.29800) showed that nine frontier LLM judges deliver only ~2.2 effective independent votes because their errors are correlated, and that the correlation is mostly not explained by item difficulty. This deposit contains a pre-registered replication on consumer hardware (six local open-weight judges, 8B to 35B parameters, four base models, 1,000 ChaosNLI-MNLI items) and a deterministic, bit-replayable instrument that converts judge logs into per-pair corroboration weights via a cross-fitted Generalised Covariance Measure over pre-registered raw-input feature extractors. All four pre-registered predictions were confirmed: effective votes 2.13 for k=6 (Kohli: 2.18 for k=9); two judge pairs known by construction to share base weights ranked first and second in error correlation, one producing identical labels on 1,000/1,000 items and receiving corroboration weight exactly 0; raw-text difficulty features explained ~0% of mean pairwise error correlation while human-annotation entropy explained 13.6% (Kohli: 13.5%). The deposit includes the technical note, the pre-registration (content hashes fixed before any judge ran), the complete instrument (pure Python standard library, no RNG), all 6,000 raw judgments with raw completions preserved, and the content-hashed analysis report. The ChaosNLI dataset is not redistributed; the deposit README explains how to regenerate the exact sample and verify its hash. The instrument is a standalone application of the measured-corroboration layer of the OFM-PBL / (In)Canon framework (UK Patent GB2606072.3): validator independence is never asserted, only measured, discounted by shared provenance, and capped below certainty.
// Source
Authors: David Antonio Tomé
Institutions: University of Stirling, Plastic Logic (United Kingdom)