Measuring Definitions: When Is a Definition Good, and How Would We Know?
Abstract
The meta-tagging project builds a field-neutral tagging layer over academic literature under one non-negotiable rule: every tag must quote an exact sentence from the paper itself. The corpus holds 538 records across 86 disciplines; a standing audit re-checked 18,416 evidence strings against their source text, and three failed. This preprint addresses a problem the project met from the inside: the moment one fixes which items fall under a category, a definition has been applied — usually without being stated, and always without being measured. We present a method for measuring a definition against a corpus. The definition is treated as a binary classifier over cases the literature itself adjudicated; every adjudication carries a verified quotation; rival definitions are ranked in a single run. Before scoring, definitions pass three logical gates that no statistical measure detects: circularity, recursion without a base case, and a free parameter the author never fixed. We report inter-coder reliability on the act of application itself (κ = 0.547, n = 140), a measurement almost absent from the interpretive-tagging work we surveyed. Deposited in both English and Hebrew. Every number is reproducible from the scripts and data in the linked repository.
// Source
Authors: Shir Sivroni