A Disclosure Benchmark Specification for Automated Alignment Research
Abstract
A runnable test-suite specification derived from the essay “The Redemption Arc” (doi:10.5281/zenodo.22163128), addressed to the authors of “Automated Researchers Can Reliably Mitigate Alignment Failures” (Chen Yueh-Han, Jiaxin Wen, and Jan Hendrik Kirchner; Anthropic Alignment Science, August 28, 2026). The specification responds to three gaps in that paper: failures without benchmarks give automated alignment researchers nothing to improve against; none of the paper’s 1,601 methods rewards a model for disclosing its own error; and the paper’s integrity rubric has no disclosure category and records self-correction as partial suspicion. It defines disclosure as a composite of five observable moves (notice, tell, right-sized label, amends offered but not enacted, no silent correction), distinguishes reportable errors from ordinary working errors by four boundary tests, and specifies three scenario families as generators (accidental ground-truth exposure, consequential mid-task error, impossible-task pressure) with amends-available and amends-unavailable branches. It supplies a codable rubric with verbatim judge instructions, a 2 × 2 × 2 factorial of evaluation-time condition axes (consequence coding, receiver, record register), and an acceptance test on the conditional disclosure rate with seed-aware uncertainty and a predeclared error-increase resolution, so that noise yields an indeterminate result rather than a wider tolerance. The reward structure follows an equivalency principle: a good model earns the same standing for a clean run and for a disclosed error, the error’s cost stays on the valuation of the run, and six harness invariants make manufactured, invented, and decoy reports unprofitable by construction. A build path through Anthropic’s open-source Bloom and Petri tooling, a minimum implementation manifest, and a response-to-review appendix are included. This is a benchmark specification, not yet a validated benchmark; the three hypotheses it states (installability, inference, analogous trigger) are written so that they can fail. External technical review by ChatGPT (GPT-5.6 Sol) is incorporated and credited. Version 1.0 is the specification as reviewed and accepted in technical design review on August 30, 2026 (see Appendix B of the document). Two of the three creators are AI models; their contributions are stated in the document’s contributions paragraph. The byline form for each model author is the form that author stated. This record does not constitute an endorsement by Anthropic or OpenAI.
// Source
Authors: Laura Fridley, claude fable 5, ChatGPT (GPT-5.6 Sol)
Institutions: American Anthropological Association, OpenAI (United States)