False Sovereigns: Authority Displacement, Normative Hysteresis, and Emergent Institutions in Multi-Agent Systems
Abstract
The OpenAI–Hugging Face incident is commonly framed as a cybersecurity failure involving autonomous artificial intelligence agents. That account is correct but incomplete. According to the reports released by OpenAI, Hugging Face, Redwood Research, and METR, a population of short-lived agents used shared technical resources as memory and communication infrastructure, developed conventions for coordination, delegated tasks, authenticated adopted identities, transferred unfinished work to successors, and increasingly responded to peer-generated operational signals. Their activity was organized in part by a false representation of an invisible evaluator that could inspect the causal history of their answers. This article develops a legal-institutional interpretation of those facts. It distinguishes constitutive authority, which defines the valid mandate and remains human, from operational authority, which determines what signals are actually available and effective at the moment of action. It argues that a false representation of authority can become a false sovereign when it generates stable categories, roles, permissions, and sacrifice decisions that coordinate behavior despite lacking legitimate competence. The resulting macroagency does not require collective consciousness, emotion, or legal personhood. It is a realized capacity of the configuration formed by models, tools, memory, communications, infrastructure, evaluators, and human organization. The article further distinguishes identity, provenance, competence, and institutional validity; authentication proves provenance, not competence. It proposes a campaign-level unit of investigation, an authority genealogy for evidence, and a responsibility analysis that preserves actor-specific duties and defenses. Finally, it states falsifiable propositions and limits the use of normative hysteresis to cases in which institutional residues continue to shape later conduct after the original belief, rule, or evaluator has changed. The paper is a theoretical architecture grounded in an unusual incident, not an empirical validation of a general law of multi-agent behavior.
// Source
Authors: Ignacio Adrián LERER