An x-ray for AI: a deterministic, model-free grounding audit of an LLM abstention benchmark
Abstract
An external, deterministic, model-free grounding audit of the Rafe and Das (2026) helmet-abstention benchmark (1,464 emergency-department injury narratives). A model-free gate re-reads each note against a sealed schema and returns a three-valued verdict; it reproduces the benchmark's separation between fabricating and abstaining models with no model in the loop, exposes nine likely gold-coding slips, and shows that a single change of prompt wording moves even a frontier model into fabrication. Version 2 corrects version 1, which contained the report only and described a harness, data, and per-cell outputs as released that were not in fact included. This version includes the evaluation harness, the Rafe and Das data as used (CC BY 4.0), and the complete per-cell outputs, and adds the adversarial negation-and-scope study the report named as its natural next step: the sealed schema's false-grounded rate under adversarial cancellation is 27 to 34 percent (development and held-out), reduced to about 5 percent by a candidate scope-hardened schema (not yet sealed), benchmarked against NegEx, ConText, and a natural-language-inference baseline; the benign-corpus false-grounded rate is zero at an effective sample of about fourteen (Wilson upper bound near 20 percent), and the nine flagged cases are gold-coding divergences. Withheld under UK patent application GB2606072.3, pending the doctoral thesis: the OFM-PBL core (the verdict-folding engine) and the sealed schema. The harness scripts import them and are provided for inspection; every reported figure is verifiable against the released per-cell outputs, but verdicts cannot be re-derived from raw notes without the withheld core and schema. See the Version 2 note in the report, DEPOSIT_README.md, and DATA.md. Version 2.1 corrects an attribution error in the negation-and-scope materials: two model-performed second-coding exercises were described as independent human coding. Both were performed by general-purpose language-model sub-agents of the same family that authored the gold labels; their perfect agreement is self-consistency, not construct validity, and the reliability framing is withdrawn. Blinding is retained; the separate three-vendor nine-case re-adjudication is unaffected; no figure or verdict changes.
// Source
Authors: David Antonio Tomé
Institutions: University of Stirling