A Blinded Multi-LLM Content-Coding Instrument: Protocol Specification and Inter-Coder Reliability Validation on an Organizational Corpus
Abstract
Content analysis in organizational research depends on inter-coder reliability, but human coding is slow and costly and single-model LLM coding is unauditable. This paper specifies a blinded multi-LLM content-coding instrument: several large language models independently code the same evidence dossiers under strict blinding, every model call is logged in structured JSONL with its full prompt and parameters, and codes are reconciled by a majority-vote-or-flag adjudication rule with a registered-before-data protocol. On a thirty-case organizational demonstration corpus the instrument reached almost-perfect inter-coder agreement (Fleiss' κ = .838). The contribution is the instrument and its reproducible reliability-validation discipline, not a substantive finding; reliability is not validity, and a held-out human-coding validation is declared future work. The complete harness, per-call logs, coded datasets, adjudication ledger, and analysis pipeline are released as the paper's computational artifact. Includes zharnikov-2026bh-multi-llm-coding-instrument.yaml (Paper Spec v0.1.0) – a machine-readable specification of the paper's claims, assumptions, and dependencies. The paper's full machine-first bundle (the SPINE claim/dependency graph and the ONTOLOGY term module) lives in the public repository; see https://github.com/spectralbranding/paper-spec for the standard. This PDF is generated programmatically from that machine-first source under a research-as-repository model.
// Source
Authors: Dmitry Zharnikov