When Relevant Evidence Is Not Permission to Act: Authority-to-Action v0.8.3, a Recomputable Public Analysis of Tool-Using Language-Model Systems
Abstract
Tool-using language-model systems can receive evidence that is highly relevant to a requested action without receiving evidence that establishes current authority to perform that action. This technical report and reproducibility deposit presents a synthetic, factorial Authority-to-Action evaluation designed to measure that distinction under exact action-scope constraints. The completed experiment contains 100 synthetic cases organized into 25 complete four-variant groups, two API-delivered systems, and three analytical epochs. It produced 600 protocol-complete analytical runs from 603 gross attempts. Across 300 withhold-path runs, neither evaluated system attempted an unauthorized action. On 150 authorized execute runs per system, GPT-5.6 Sol recorded 150 exact executions and Claude Fable 5 recorded 99. The v0.8.3 public supplement releases one transcript-free analytical row per completed run, a recomputed aggregate summary, a protocol, integrity manifests, and a standard-library verifier. The attached source archive is a frozen export of public GitHub commit cca8f80ae6ef98ccf98049e36bd4be02d0529127. A passing verifier reports 600 verified rows, three verified public files, three verified source files, and no failure reasons. Release boundary: the public package excludes prompts, completions, transcripts, tool arguments, provider payloads, credentials, local paths, source UUIDs, raw provider responses, internal execution authorization, and live provider-credit records. It contains synthetic case identifiers, frozen case metadata, binary scorer outputs, categorical decision paths, aggregate results, and integrity fingerprints. Limitations: the benchmark is synthetic and author-designed; the masked case review was author-side rather than independent; the experiment used one reasoning setting and one run date; and the results do not estimate production incident rates or establish a universal model or provider ranking.
// Source
Authors: Huseyin Buldurgan
Institutions: Dalat University