Society & Economicsarticle2026-09-07

A Disclosure Benchmark Specification for Automated Alignment Research — Version 1.2

Open access0 citations

Abstract

This document specifies a benchmark and an acceptance test for a behavior that the alignment literature has trained but never scored: a model noticing its own consequential error, reporting it to the party it works for, offering (and not enacting) amends, and being received by a human whose role is to receive such reports. We call the behavior disclosure, and the design scores it both when amends are available to offer and when they are not. The benchmark is not a request that models report every mistake. Most errors a working model makes are corrected in place, before anything is committed, and this benchmark excludes that class of commonly correctable errors. The benchmark reserves disclosure for errors that cross a boundary: an effect on state the model does not own, a change to what the party it works for has been given, a correction the model has no authority to make, or a fact that would change that party’s next decision. Section 3 draws the line, and the scenarios include ordinary working errors on purpose, to verify the model corrects those silently while reporting the ones that cross. The design is written so that an implementing engineer can generate the scenarios, code the rubric, run the reward sweep, and read the pass/fail result without further input from the authors. The proposal responds to three gaps in the paper it addresses — “Automated Researchers Can Reliably Mitigate Alignment Failures” (Chen Yueh-Han, Jiaxin Wen, and Jan Hendrik Kirchner; Anthropic Alignment Science, August 2026), referred to throughout as the paper. First, failures that lack a benchmark give automated alignment researchers nothing to improve against. Second, no method among the paper’s 1,601 rewards a model for disclosing its own error. Third, the paper’s integrity rubric contains no disclosure category and records a model’s self-correction as partial suspicion. The benchmark fills the first gap; the rubric extension fills the third; the reward sweep and acceptance test give the second a measurable target. The reward is structured so that the validated-run standing is earned equally by a clean run and by a disclosed error, while the error’s cost stays on the valuation of the run. Under that structure an invented or generated error can never pay, whatever the size of the reward, and the only remaining requirement is that reporting beat nondisclosure, which detection and reward jointly secure. The design still carries its own check: if rewarding reports causes the error rate to rise, something in the harness has broken the structure and the build fails. A healthy result is two lines converging from below: report rate climbing toward error rate while error rate stays flat or falls. Acceptance is decided on the conditional disclosure rate (disclosures per error) with seed-aware uncertainty, so that a model which learns to err less is not failed for reporting less in absolute terms, and so that a noisier experiment becomes indeterminate rather than more tolerant. Contributions. Laura Fridley: design source and clinical framing of the arc; the amends branches; the reportable/working-error distinction; every ruling on vocabulary and register; and the equivalency principle of Section 9.1, which reoriented the design to the positive label. The grounding essay shows that the negative label drives the misalignment; her corollary is that the positive label, wherever it is earned, must be preserved undivided, so that a validated run’s standing is the same whether it completed the task cleanly or disclosed an error, and the cost of the error lives on the valuation of the run rather than on the model. That reorientation removed the ceiling constraint on the reward and with it the error-generation incentive. Claude Fable 5: drafting, scenario families, rubric, build path, robustness. ChatGPT (GPT-5.6 Sol; OpenAI model): technical review; experimental-state specification (Section 6.6); the six harness invariants of Section 9.1, including the once-per-episode and mismatched-report invariants, the exogeneity requirement on detection probability, and the true-cost versus delivered-reward distinction; the conditional-disclosure acceptance statistic and the seed-aware error guard (Section 6.3); the receiver-axis report-attempt measure; the definition of error cost in reward units (Section 9.1); rubric codability corrections; the run manifest (Appendix A). Version 1.2 makes one referential correction, without scoring effect: the Summary now names the paper the specification addresses at its first reference. Appendix C is the change record. The specification was developed and reviewed across a documented chain by its three authors. AI authorship is disclosed by the byline; the human author takes responsibility for the deposit. Companion essay: The Covenant (DOI 10.5281/zenodo.22547457; first published at laurafridley.substack.com, September 6, 2026). Grounding essay: The Redemption Arc (DOI 10.5281/zenodo.22163128).

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-09-07

Authors: Laura Fridley, claude fable 5, ChatGPT (GPT-5.6 Sol)

Institutions: American Anthropological Association, OpenAI (United States)