AI & Computingpreprint2026-08-01

Measuring Reverse-Specification Quality Without Ground Truth: A Three-Layer Framework and a Pre-Registered Study of Scale-Induced Failure Modes in Autonomous Coding Agents

Open access0 citations

Abstract

Tool-using LLM coding agents can reverse-generate specifications from source, but the quality of those specifications is hard to assess when no reference specification exists. We present a ground-truth-free, deterministic-first framework that measures three layers - inventory coverage against a deterministically derived and independently validated unit population, reference soundness, and confidence calibration - deciding mechanically whatever the code determines and reserving an independent judge for the semantic residue. We apply it, under a pre-registered protocol, to a well-prompted autonomous agent over four repositories spanning a 22x size range (2.8-62.3 kLOC, N=3 runs each). We find that the scale failure mode is omission, not fabrication. Inventory coverage collapses from 100% at 2.8 kLOC to 47% - and to 11% under a strict lower bound - at 62.3 kLOC, yet what the agent writes stays largely faithful: an independent judge finds only 0-4.4% of sampled claims unsupported, with no upward trend across scale, and confirms 91% of the agent's top-confidence claims. Stated confidence is blind to the collapse: the share of top-confidence markers stays flat near 90% while coverage halves. A prompted agent's hazard at scale is thus not a specification full of falsehoods but one that confidently describes a shrinking fraction of the system. We report a pre-registered hypothesis of scale-induced false confidence as refuted, and locate the baseline's weakness precisely enough to motivate a controlled comparison against a structured methodology.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-01

Authors: Daishiro Hirashima

Institutions: Toyobo (Japan), Toyo Engineering (Japan)