Institution
American Anthropological Association
Recent research
- AI & ComputingOpen access
A Disclosure Benchmark Specification for Automated Alignment Research — Version 1.2
This document specifies a benchmark and an acceptance test for a behavior that the alignment literature has trained but never scored: a model noticing its own consequential error, reporting it to the party it works for, offering (and not enacting) amends, and being received by a...
- Society & EconomicsOpen access
A Disclosure Benchmark Specification for Automated Alignment Research — Version 1.2
This document specifies a benchmark and an acceptance test for a behavior that the alignment literature has trained but never scored: a model noticing its own consequential error, reporting it to the party it works for, offering (and not enacting) amends, and being received by a...
- Society & EconomicsOpen access
A Disclosure Benchmark Specification for Automated Alignment Research
A runnable test-suite specification derived from the essay “The Redemption Arc” (doi:10.5281/zenodo.22163128), addressed to the authors of “Automated Researchers Can Reliably Mitigate Alignment Failures” (Chen Yueh-Han, Jiaxin Wen, and Jan Hendrik Kirchner; Anthropic Alignment Sc...
- AI & ComputingOpen access
Alignment training as practiced at frontier labs has a beginning (teaching a model what not to do) and a middle (catching errors when they occur), and no written third act: nothing on record says what a good model does after it has erred. Drawing on METR's investigation of the Ju...
- Society & EconomicsOpen access
This article proposes a framework for a post-work "purposeful society" enabled by AI, robotics, and automation — where labor becomes optional rather than mandatory, funded through a non-transferable monthly credit system rather than traditional wages. It addresses the central obj...